Repository navigation
Request: Support SME for all GEMM kernels on Apple silicon #5841
Description
Activity
SME GEMM support is currently hampered by the interdependence between GEMM,SYMM and (to a lesser degree) TRMM kernels in the code structure inherited from GotoBLAS. (Worked around in the ARM-contributed PR #5573 but I have been hesitant to merge it - perhaps I should, as I haven't had time to develop the alternative within the existing framework).
Currently there is no research group and not much of a permanent team behind OpenBLAS, so progress largely depends on contributed code.
That said, I wonder how your implementation of SGEMM_NN compares with the one M4/M5 fast path (bypassing the regular blocked GEMM algorithm) the OpenBLAS codebase hasThank you for this proposal! When I pull this implementation into my BLAS benchmark, I do achieve about 400 GFLOPS with DGEMM_NN on M4 chip.
Small Matrix
config: gemm = 32 x 32 times 32 x 32 repeat = 32 warmup = 8 kernel impl seconds GFLOP/s error ======================================================================== gemm Reference 1.050000e-05 6.24 0.000000e+00 gemm stdBLAS 2.495800e-05 2.63 0.000000e+00 gemm OpenBLAS 3.708000e-06 17.67 0.000000e+00 gemm Accelerate 8.750000e-07 74.90 0.000000e+00 gemm ARMv9.2 2.208000e-06 29.68 0.000000e+00 ------------------------------------------------------------------------
Large Matrix
config: gemm = 512 x 512 times 512 x 512 repeat = 32 warmup = 8 kernel impl seconds GFLOP/s error ======================================================================== gemm Reference 1.634788e-02 16.42 0.000000e+00 gemm stdBLAS 1.182299e-01 2.27 0.000000e+00 gemm OpenBLAS 4.763750e-03 56.35 0.000000e+00 gemm Accelerate 5.980420e-04 448.86 0.000000e+00 gemm ARMv9.2 6.580830e-04 407.91 3.449122e+01 ------------------------------------------------------------------------
Experimenting with your code added as another fast path (bypassing Goto's block matrix algorithm and its multithreading), I'm seeing about doubled performance in SGEMM; and around 10 percent speedup over the (block matrix) NEON kernel in DGEMM on my M4 mini. (ISTR the SME hardware is single-precision at least on the smaller models ?)
Notes:- this is with the omp-parallel loops serialized for simplicity, and the C++ trivially converted to C
- I needed to add the dummy assembler routine from OpenBLAS' current "small-matrix SME" kernels that only marks everything as clobbered,
otherwise any subsequent floating point operation in the code would become totally erratic - the real number versions of your codes pass the BLAS tests, but unfortunately the complex versions do not, failing already for a M,N,K=1 case with alpha=(1,0) and beta zero (also reproducible when linking a stripped-down cblat3.f against your unmodified libsme_gemm.a)
fixed my issues with the complex codes - mostly they needed trivial additions for the conjugate and conjugate transpose cases, all tests passing now, will benchmark tomorrow
ZGEMM speedup is about 15 percent on M4mini for problem sizes around 2000x2000 (320 GFlops vs 280), I still need to run the comparison for CGEMM (tried to go all the way to 5000x5000 with ZGEMM, which may not have been such a bright idea). Will update tomorrow or so, including performance for 32x32 and 512x512 to compare with yhu1729's data
Hmm, I'm currently seeing a slowdown for the 32x32 case with my adaption of vlovero's code vs. the stock OpenBLAS NEON kernel, 23 GFlops vs 47 previously, while at 512x512 the new kernel outperforms the old one (263 GFlops vs 181 in the GEMM benchmark.
In SGEMM, it is even more extreme, only 19GFlops instead of 79 with theexisting "small matrix" SMENEON kernel at 32x32, but 505 vs 340 at 512x512. As noted above, this is without the OMP parallelization of the setup loops in the original code. Figures are broadly similar for the complex kernels - those will probably gain a few Flops in cleanup.
Not sure where the discrepancy between my results and yhu1729's benchmark comes from, but my M4mini may be thermally constrained (or there's still something wrong with the GFlops calculation in the stock gemm benchmark even after correcting the cputime code that was returning too low performance numbers on Macs)Reacted by Yifan HuUpdated benchmark on M4.
Compiled on macOS 26.6.1 with
clang 22.1.8andgfortran 16.1.0. Benchmark context:Run on (10 X 24 MHz CPU s) CPU Caches: L1 Data 64 KiB L1 Instruction 128 KiB L2 Unified 4096 KiB (x10)
vlovero's kernels at vlovero/ARMv9.2-gemm@b0b91d4
Benchmark Time CPU Iterations Byte FLOP ----------------------------------------------------------------------------------------------------- gemm/s/vlovero_arm/m:32/n:32/k:32 5.61 us 5.50 us 32 2.23418G/s 11.9156G/s gemm/d/vlovero_arm/m:32/n:32/k:32 6.06 us 5.88 us 32 4.18315G/s 11.1551G/s gemm/s/vlovero_arm/m:128/n:128/k:128 18.8 us 18.8 us 32 10.4683G/s 223.324G/s gemm/d/vlovero_arm/m:128/n:128/k:128 35.6 us 35.5 us 32 11.0668G/s 118.045G/s gemm/s/vlovero_arm/m:512/n:512/k:512 337 us 337 us 32 9.33277G/s 796.397G/s gemm/d/vlovero_arm/m:512/n:512/k:512 930 us 929 us 32 6.77024G/s 288.864G/s
OpenBLAS v0.3.34
Benchmark Time CPU Iterations Byte FLOP ----------------------------------------------------------------------------------------------------- gemm/s/openblas/m:32/n:32/k:32 5.39 us 5.22 us 32 2.35459G/s 12.5578G/s gemm/d/openblas/m:32/n:32/k:32 6.01 us 5.84 us 32 4.20552G/s 11.2147G/s gemm/s/openblas/m:128/n:128/k:128 34.2 us 33.9 us 32 5.79858G/s 123.703G/s gemm/d/openblas/m:128/n:128/k:128 47.4 us 46.5 us 32 8.46194G/s 90.2607G/s gemm/s/openblas/m:512/n:512/k:512 788 us 786 us 32 4.00156G/s 341.467G/s gemm/d/openblas/m:512/n:512/k:512 1503 us 1502 us 32 4.18881G/s 178.722G/s
ArmPL v26.07
Benchmark Time CPU Iterations Byte FLOP ----------------------------------------------------------------------------------------------------- gemm/s/armpl/m:32/n:32/k:32 6.18 us 6.03 us 32 2.03739G/s 10.8661G/s gemm/d/armpl/m:32/n:32/k:32 6.66 us 6.66 us 32 3.69217G/s 9.84578G/s gemm/s/armpl/m:128/n:128/k:128 33.6 us 33.7 us 32 5.84165G/s 124.622G/s gemm/d/armpl/m:128/n:128/k:128 34.1 us 33.9 us 32 11.6079G/s 123.817G/s gemm/s/armpl/m:512/n:512/k:512 579 us 579 us 32 5.4345G/s 463.744G/s gemm/d/armpl/m:512/n:512/k:512 892 us 892 us 32 7.05221G/s 300.894G/s
Accelerate v1.11
Benchmark Time CPU Iterations Byte FLOP ----------------------------------------------------------------------------------------------------- gemm/s/accelerate/m:32/n:32/k:32 4.76 us 4.84 us 32 2.53688G/s 13.53G/s gemm/d/accelerate/m:32/n:32/k:32 5.24 us 5.31 us 32 4.62607G/s 12.3362G/s gemm/s/accelerate/m:128/n:128/k:128 13.7 us 13.7 us 32 14.3641G/s 306.433G/s gemm/d/accelerate/m:128/n:128/k:128 28.7 us 28.9 us 32 13.6179G/s 145.257G/s gemm/s/accelerate/m:512/n:512/k:512 249 us 249 us 32 12.6255G/s 1.07738T/s gemm/d/accelerate/m:512/n:512/k:512 853 us 853 us 32 7.37947G/s 314.857G/s
Thanks for the benchmark updates. In my latest tests, vlovero's kernels outclass the old NEON ones in single-threaded runs, it is only in multithreaded where the advantage of their single-threaded fast path over the classic blocked-GEMM algorithm is less pronounced (but also thermal throttling of my M4mini and the known issues with buffer assignment bottlenecks under load complicate the picture).
I have put what I currently have into PR #5971 - @vlovero not sure if getting your code half-blindly plugged into OpenBLAS even was what you intended, but I cannot compete with you in this field, and you're credited in the Contributors list. (Also feel free to copy the PR and resubmit so that you're properly in the git history, if you agree with this at all)Reacted by Yifan Hu@martin-frbg Everything seems good to me. Before a PR is merged I need to figure out a compiler dependent bug I'm experiencing. AppleClang seems to be optimizing away code blocks that handle the zero padding only when the size of the matrices are larger than the size of the static packing matrices. I've recompiled it with gcc and do not have the same issues. The sanitizers aren't catching anything and no compiler warnings so I'm not sure if it's UB or an actual compiler bug.
@martin-frbg I was able to get all of the tests to work on AppleClang with aggressive compiler optimizations enabled by using much more aggressive register clobbering. I'm not sure which register was the culprit (there's so many), but I think one of them may have been leading to some type of UB I don't quite understand.
The updated code is on my repo now. I can duplicate your PR or do whatever is easiest if you think things are ready for OpenBLAS
Thank you - I'll just add your changes to the clobber lists as I revamp the PR to include support for DYNAMIC_ARCH builds.
(I've just tried with SGEMM, and indeed the need for the additional "clobber list only" inline assembly routines is resolved with your change.)
One very minor issue that still bothers me is that the Fortran test codes (test/?blat3) complain that all of IEEE_INVALID_FLAG IEEE_DIVIDE_BY_ZERO IEEE_OVERFLOW_FLAG IEEE_UNDERFLOW_FLAG are signalling, suggesting that the fpsr got trashed by SME.
Sneaking in an"msr fpsr, xzr\n\t"after thesmstoptakes care of it, but I'm not sure if that is a sane way to address it ? (I'm fairly certain that these condition flags are spurious, of course, but I wouldn't want to suppress them at the gfortran option level)I think a manual clearing of the flags may be necessary; looking at ARM documentation says:
When the Effective value of PSTATE.SM is changed by any method from 0 to 1 or from 1 to 0, the FPSR is set to the value 0x0000_0000_0800_009f, setting all the cumulative status bits to 1.
And when I manually step through the assembly using lldb, every time an SME instruction gets called, the fpsr register gets changed back to 0x800009f (even if I set it to zero after entering streaming mode).
I'll look into this more to see if there is a way to detect IEEE exceptions when using streaming.
Going through the ARM documentation it seems like it is not possible using the fpsr register because they are purposely not updated to maximize throughput
RTGSKG: Produce the expected IEEE 754 default result but do not update the FPSR cumulative exception flag bits.
So zeroing it out seems like the way to go to suppress the fortran test code complaints.
Performance was lacking compared to Accelerate and ArmPL for complex arithmetic. I changed to packing dimensions and got 8-20% speedups.
On Apple silicon (I'm using base M5), gemm gets poor performance relative to Apple's Accelerate framework. In my benchmarks, dgemm achieves about 64 GFLOPS at its peak (single thread) while Accelerate is getting about 480 GLOPS.
Playing around with the SME instructions, I was able to get about 440 GFLOPS for dgemm on a base M5. I've also written kernels for sgemm, cgemm, and zgemm that are all in my repo here. The code is not optimized for small matrices yet but it would be nice to get more performance out of the hardware.