Repository navigation
Add NEON for cwp_convolve_x, cwp_convolve_y, cwp_convolve_2d_copy - #5508
jjustiss-apple wants to merge 2 commits into
Conversation
| // constant 3 or 5 and BITS is 0. Falls back to the original 4-op sequence | ||
| // for non-standard values. | ||
| // Captures round_const, round_shift, bits_shift from enclosing scope. | ||
| #define CWP_X_POST_FILTER(sum, ROUND0, BITS, offset) \ |
There was a problem hiding this comment.
All the callsites to this macro have BITS as 0 (also only called by other macros)
(BITS) == 4and(BITS) == 3branches are dead code.- at compile time, the branch
else if ((BITS) > 0)will be ignored. I'm not sure what the consequences are from this... ifBITSwill always be 0, it can be removed from this macro.
There was a problem hiding this comment.
You're right it's all dead code (now removed). The compiler was removing it which is why it didn't affect the speed.
I found that I can run clang -E to unfold the macros and then use regex to find literal comparisons, which flagged all these.
A challenge I'm learning with macros is that -Wunreachable-code is suppressed in clang, which requires extra handling.
Thank you for flagging this!
| ConvolveParams *conv_params, int bd) { | ||
| const int tap_y = get_filter_tap(filter_params_y, subpel_y_qn); | ||
|
|
||
| (void)tap_y; |
| s8 = s9; \ | ||
| s9 = s10; \ | ||
| s10 = s11; \ | ||
| height--; \ |
There was a problem hiding this comment.
we can unroll 4 rows each iteration here (same for CONV_Y_12TAP_8)
4b04cd0 to
ef68e0e
Compare
|
@jianj-g everything should be addressed, thanks again for your thorough review. I'm tightening my own self-review processes and hopefully will have fewer of these dead code issues going forward. |
Adds optimized NEON intrinsics for av2_highbd_cwp_convolve_x,
av2_highbd_cwp_convolve_y, and av2_highbd_cwp_convolve_2d_copy.
Also improves the existing 2D path (landed in #5425) with symmetric
coefficient folding and MODE/RBITS dispatch.
Optimization techniques: immediate-shift post-filter, symmetric
coefficient folding (6-tap), MODE/RBITS compile-time dispatch,
s32 halving add for compound avg, tap-count specialized kernels
(2/4/6/8/12-tap horizontal, 4/6/8-tap vertical).
Micro-benchmarks (Apple M2 Ultra P-core):
CTC Results (RA, cpu-used=1, 33 frames, A5):