Repository navigation
Remove redundant operations in the encoder - #5511
Merged
leolzhao merged 10 commits intoOct 8, 2026
Merged
Conversation
av2_ml_part_split_features_square() and av2_ml_part_split_features_none() clear the whole intrapred buffer for every intra mode already. av2_predict_intra_block() always writes the full transform block, and the variance function only reads those blocks, so the memset redundant. Bit-exact change.
warehouse_efficients_txb() rescanned all coefficients to count nonzeros, but the count is only used when parity hiding is on. So, we skip the counting when parity hiding is disabled (e.g. always for chroma). Bit-exact change.
av2_decide_states_c() and av2_decide_states_q1_c() clear `rdCost` and `rdCost_zero` for every coefficient position. In both functions, every entry that is read is written first, so the memsets were redundant. Bit-exact change, and affects C-only build only.
av2_cdef_search() was copying the same filter-block window from the reconstruction into the CDEF input buffer for each of up to 64 strength candidates. The source window does not depend on the strength, and av2_cdef_filter_fb() only reads the input buffer, so copy it once per filter block and plane. Bit-exact change.
In handle_smooth_inter_intra_mode(), the 'recompute prediction if required' condition was always true: interintra_mode_reuse was computed after the search, when *best_interintra_mode is always a valid mode. So the intra predictor and the interintra blend were always rebuilt, even when the best mode was the last one evaluated and its prediction was already in place. Track the last mode built by the search and rebuild only if the best mode differs from it. Bit-exact change.
fill_dv_costs() built the DV shell cost tables (up to 32K entries for 1/8-pel) for all 7 MV precisions, but IntraBC only uses the precisions in av2_intraBc_precision_sets (ONE_PEL and QTR_PEL). Build and copy only those. Add an assert in get_vq_mvd_cost() that IntraBC costs are only queried for precisions in that set. Bit-exact change.
- av2_cost_color_map() / av2_tokenize_color_map(): direction 1 is only allowed when both plane dimensions are < 64. Skip the direction-1 cost pass (and in tokenize, both cost passes) when it cannot be selected. - cost_and_tokenize_map(): in rate-only mode, derive the per-pixel color index context (which scores neighbors and sorts the color order) only for pixels whose color index is actually coded, i.e. not in copied lines or in non-first pixels of identity rows. Bit-exact change.
search_filter_offsets() always copied the plane into last_frame_uf and ran a full-plane deblock + SSE at the starting offset. In three cases that configuration was already evaluated by the previous search, and last_frame_uf already holds the unfiltered plane: - luma VERT_EDGE search: starts from the single-offset (both edges) result, whose SSE is best_single_sse; - luma HORZ_EDGE search: starts from the VERT_EDGE search result, whose SSE is best_dual_sse; - chroma pass 1: starts from the pass 0 result, whose SSE is best_filter_sse[plane]. Pass the known SSE in and skip the copy and the initial filter run. Bit-exact change.
compute_stats_for_wienerns_filter() looped over the classes on the outside, scanning the whole restoration unit and looking up every pixel's class once per class (16 times for luma), while each pixel only contributes to its own class. Do a single raster pass and accumulate each pixel into its class. The accumulation order within each class is the same raster order as before, so the (floating-point) statistics are identical. Bit-exact change.
get_rate_dist_def_luma_avx2_impl() and get_rate_dist_lf_luma_avx2_impl() stored the base rates to rd->rate / rd->rate_eob, then reloaded, added the mid-range and high-range (golomb) costs, and stored them again, up to three times per coefficient. The intrinsic loads and stores may alias, so the compiler could not remove these round trips even in the fused av2_trellis_loop_diagonal_st8_avx2() loop. Keep the rates in registers and store them once. The rare scalar golomb fallback becomes an equivalent vector add. The DC path of the LF kernel, which adds per-state costs in scalar code, stores first as before. Bit-exact change.
urvangjoshi
force-pushed
the
mg__remove_redundant_ops
branch
from
October 8, 2026 19:02
f2bab4d to
f5b5f32
Compare
leolzhao
approved these changes
Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Note: This is a series of independent commits -- all are bit-exact.
Pls see description of each commit and review / merge them separately.
Bit-exact results with ~0.4% speed-up at speed 1.