Skip to content

Remove redundant operations in the encoder - #5511

Merged
leolzhao merged 10 commits into
AOMediaCodec:av2-encfrom
urvangjoshi:mg__remove_redundant_ops
Oct 8, 2026
Merged

leolzhao merged 10 commits into
AOMediaCodec:av2-encfrom
urvangjoshi:mg__remove_redundant_ops

Conversation

@urvangjoshi

@urvangjoshi urvangjoshi commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Note: This is a series of independent commits -- all are bit-exact.
Pls see description of each commit and review / merge them separately.

Bit-exact results with ~0.4% speed-up at speed 1.

av2_ml_part_split_features_square() and av2_ml_part_split_features_none()
clear the whole intrapred buffer for every intra mode already.
av2_predict_intra_block() always writes the full transform block, and the
variance function only reads those blocks, so the memset redundant.

Bit-exact change.
warehouse_efficients_txb() rescanned all coefficients to count nonzeros,
but the count is only used when parity hiding is on.

So, we skip the counting when parity hiding is disabled (e.g. always for chroma).

Bit-exact change.
av2_decide_states_c() and av2_decide_states_q1_c() clear `rdCost` and
`rdCost_zero` for every coefficient position. In both functions, every
entry that is read is written first, so the memsets were redundant.

Bit-exact change, and affects C-only build only.
av2_cdef_search() was copying the same filter-block window from the
reconstruction into the CDEF input buffer for each of up to 64 strength
candidates. The source window does not depend on the strength, and
av2_cdef_filter_fb() only reads the input buffer, so copy it once per
filter block and plane.

Bit-exact change.
In handle_smooth_inter_intra_mode(), the 'recompute prediction if
required' condition was always true: interintra_mode_reuse was computed
after the search, when *best_interintra_mode is always a valid mode. So
the intra predictor and the interintra blend were always rebuilt, even
when the best mode was the last one evaluated and its prediction was
already in place.

Track the last mode built by the search and rebuild only if the best
mode differs from it.

Bit-exact change.
fill_dv_costs() built the DV shell cost tables (up to 32K entries for
1/8-pel) for all 7 MV precisions, but IntraBC only uses the precisions
in av2_intraBc_precision_sets (ONE_PEL and QTR_PEL). Build and copy only
those. Add an assert in get_vq_mvd_cost() that IntraBC costs are only
queried for precisions in that set.

Bit-exact change.
- av2_cost_color_map() / av2_tokenize_color_map(): direction 1 is only
  allowed when both plane dimensions are < 64. Skip the direction-1 cost
  pass (and in tokenize, both cost passes) when it cannot be selected.
- cost_and_tokenize_map(): in rate-only mode, derive the per-pixel color
  index context (which scores neighbors and sorts the color order) only
  for pixels whose color index is actually coded, i.e. not in copied
  lines or in non-first pixels of identity rows.

Bit-exact change.
search_filter_offsets() always copied the plane into last_frame_uf and
ran a full-plane deblock + SSE at the starting offset. In three cases
that configuration was already evaluated by the previous search, and
last_frame_uf already holds the unfiltered plane:
- luma VERT_EDGE search: starts from the single-offset (both edges)
  result, whose SSE is best_single_sse;
- luma HORZ_EDGE search: starts from the VERT_EDGE search result, whose
  SSE is best_dual_sse;
- chroma pass 1: starts from the pass 0 result, whose SSE is
  best_filter_sse[plane].

Pass the known SSE in and skip the copy and the initial filter run.

Bit-exact change.
compute_stats_for_wienerns_filter() looped over the classes on the
outside, scanning the whole restoration unit and looking up every
pixel's class once per class (16 times for luma), while each pixel only
contributes to its own class.

Do a single raster pass and accumulate each pixel into its class. The
accumulation order within each class is the same raster order as
before, so the (floating-point) statistics are identical.

Bit-exact change.
get_rate_dist_def_luma_avx2_impl() and get_rate_dist_lf_luma_avx2_impl()
stored the base rates to rd->rate / rd->rate_eob, then reloaded, added
the mid-range and high-range (golomb) costs, and stored them again, up
to three times per coefficient. The intrinsic loads and stores may
alias, so the compiler could not remove these round trips even in the
fused av2_trellis_loop_diagonal_st8_avx2() loop.

Keep the rates in registers and store them once. The rare scalar golomb
fallback becomes an equivalent vector add. The DC path of the LF kernel,
which adds per-state costs in scalar code, stores first as before.

Bit-exact change.
@leolzhao
leolzhao merged commit 8ebdc75 into AOMediaCodec:av2-enc Oct 8, 2026
107 checks passed
@urvangjoshi
urvangjoshi deleted the mg__remove_redundant_ops branch October 8, 2026 23:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants