Skip to content

Request: Support SME for all GEMM kernels on Apple silicon #5841

Description

@vlovero

On Apple silicon (I'm using base M5), gemm gets poor performance relative to Apple's Accelerate framework. In my benchmarks, dgemm achieves about 64 GFLOPS at its peak (single thread) while Accelerate is getting about 480 GLOPS.

Playing around with the SME instructions, I was able to get about 440 GFLOPS for dgemm on a base M5. I've also written kernels for sgemm, cgemm, and zgemm that are all in my repo here. The code is not optimized for small matrices yet but it would be nice to get more performance out of the hardware.

Activity

  1. martin-frbg commented on Jun 17, 2026

    @martin-frbg
    Collaborator

    SME GEMM support is currently hampered by the interdependence between GEMM,SYMM and (to a lesser degree) TRMM kernels in the code structure inherited from GotoBLAS. (Worked around in the ARM-contributed PR #5573 but I have been hesitant to merge it - perhaps I should, as I haven't had time to develop the alternative within the existing framework).
    Currently there is no research group and not much of a permanent team behind OpenBLAS, so progress largely depends on contributed code.
    That said, I wonder how your implementation of SGEMM_NN compares with the one M4/M5 fast path (bypassing the regular blocked GEMM algorithm) the OpenBLAS codebase has

  2. yhu1729 commented on Jun 19, 2026

    @yhu1729

    Thank you for this proposal! When I pull this implementation into my BLAS benchmark, I do achieve about 400 GFLOPS with DGEMM_NN on M4 chip.

    Small Matrix

    config:
      gemm   = 32 x 32 times 32 x 32
      repeat = 32
      warmup = 8
    
    kernel    impl                 seconds        GFLOP/s              error
    ========================================================================
    gemm      Reference       1.050000e-05           6.24       0.000000e+00
    gemm      stdBLAS         2.495800e-05           2.63       0.000000e+00
    gemm      OpenBLAS        3.708000e-06          17.67       0.000000e+00
    gemm      Accelerate      8.750000e-07          74.90       0.000000e+00
    gemm      ARMv9.2         2.208000e-06          29.68       0.000000e+00
    ------------------------------------------------------------------------

    Large Matrix

    config:
      gemm   = 512 x 512 times 512 x 512
      repeat = 32
      warmup = 8
    
    kernel    impl                 seconds        GFLOP/s              error
    ========================================================================
    gemm      Reference       1.634788e-02          16.42       0.000000e+00
    gemm      stdBLAS         1.182299e-01           2.27       0.000000e+00
    gemm      OpenBLAS        4.763750e-03          56.35       0.000000e+00
    gemm      Accelerate      5.980420e-04         448.86       0.000000e+00
    gemm      ARMv9.2         6.580830e-04         407.91       3.449122e+01
    ------------------------------------------------------------------------
  3. martin-frbg commented on Jul 30, 2026

    @martin-frbg
    Collaborator

    Experimenting with your code added as another fast path (bypassing Goto's block matrix algorithm and its multithreading), I'm seeing about doubled performance in SGEMM; and around 10 percent speedup over the (block matrix) NEON kernel in DGEMM on my M4 mini. (ISTR the SME hardware is single-precision at least on the smaller models ?)
    Notes:

    • this is with the omp-parallel loops serialized for simplicity, and the C++ trivially converted to C
    • I needed to add the dummy assembler routine from OpenBLAS' current "small-matrix SME" kernels that only marks everything as clobbered,
      otherwise any subsequent floating point operation in the code would become totally erratic
    • the real number versions of your codes pass the BLAS tests, but unfortunately the complex versions do not, failing already for a M,N,K=1 case with alpha=(1,0) and beta zero (also reproducible when linking a stripped-down cblat3.f against your unmodified libsme_gemm.a)
  4. martin-frbg commented on Aug 4, 2026

    @martin-frbg
    Collaborator

    fixed my issues with the complex codes - mostly they needed trivial additions for the conjugate and conjugate transpose cases, all tests passing now, will benchmark tomorrow

  5. martin-frbg commented on Aug 7, 2026

    @martin-frbg
    Collaborator

    ZGEMM speedup is about 15 percent on M4mini for problem sizes around 2000x2000 (320 GFlops vs 280), I still need to run the comparison for CGEMM (tried to go all the way to 5000x5000 with ZGEMM, which may not have been such a bright idea). Will update tomorrow or so, including performance for 32x32 and 512x512 to compare with yhu1729's data

  6. martin-frbg commented on Aug 8, 2026

    @martin-frbg
    Collaborator

    Hmm, I'm currently seeing a slowdown for the 32x32 case with my adaption of vlovero's code vs. the stock OpenBLAS NEON kernel, 23 GFlops vs 47 previously, while at 512x512 the new kernel outperforms the old one (263 GFlops vs 181 in the GEMM benchmark.
    In SGEMM, it is even more extreme, only 19GFlops instead of 79 with the existing "small matrix" SME NEON kernel at 32x32, but 505 vs 340 at 512x512. As noted above, this is without the OMP parallelization of the setup loops in the original code. Figures are broadly similar for the complex kernels - those will probably gain a few Flops in cleanup.
    Not sure where the discrepancy between my results and yhu1729's benchmark comes from, but my M4mini may be thermally constrained (or there's still something wrong with the GFlops calculation in the stock gemm benchmark even after correcting the cputime code that was returning too low performance numbers on Macs)

  7. yhu1729 commented on Aug 10, 2026

    @yhu1729

    Updated benchmark on M4.

    Compiled on macOS 26.6.1 with clang 22.1.8 and gfortran 16.1.0. Benchmark context:

    Run on (10 X 24 MHz CPU s)
    CPU Caches:
      L1 Data 64 KiB
      L1 Instruction 128 KiB
      L2 Unified 4096 KiB (x10)

    vlovero's kernels at vlovero/ARMv9.2-gemm@b0b91d4

    Benchmark                                     Time             CPU   Iterations       Byte       FLOP
    -----------------------------------------------------------------------------------------------------
    gemm/s/vlovero_arm/m:32/n:32/k:32          5.61 us         5.50 us           32 2.23418G/s 11.9156G/s
    gemm/d/vlovero_arm/m:32/n:32/k:32          6.06 us         5.88 us           32 4.18315G/s 11.1551G/s
    gemm/s/vlovero_arm/m:128/n:128/k:128       18.8 us         18.8 us           32 10.4683G/s 223.324G/s
    gemm/d/vlovero_arm/m:128/n:128/k:128       35.6 us         35.5 us           32 11.0668G/s 118.045G/s
    gemm/s/vlovero_arm/m:512/n:512/k:512        337 us          337 us           32 9.33277G/s 796.397G/s
    gemm/d/vlovero_arm/m:512/n:512/k:512        930 us          929 us           32 6.77024G/s 288.864G/s

    OpenBLAS v0.3.34

    Benchmark                                     Time             CPU   Iterations       Byte       FLOP
    -----------------------------------------------------------------------------------------------------
    gemm/s/openblas/m:32/n:32/k:32             5.39 us         5.22 us           32 2.35459G/s 12.5578G/s
    gemm/d/openblas/m:32/n:32/k:32             6.01 us         5.84 us           32 4.20552G/s 11.2147G/s
    gemm/s/openblas/m:128/n:128/k:128          34.2 us         33.9 us           32 5.79858G/s 123.703G/s
    gemm/d/openblas/m:128/n:128/k:128          47.4 us         46.5 us           32 8.46194G/s 90.2607G/s
    gemm/s/openblas/m:512/n:512/k:512           788 us          786 us           32 4.00156G/s 341.467G/s
    gemm/d/openblas/m:512/n:512/k:512          1503 us         1502 us           32 4.18881G/s 178.722G/s

    ArmPL v26.07

    Benchmark                                     Time             CPU   Iterations       Byte       FLOP
    -----------------------------------------------------------------------------------------------------
    gemm/s/armpl/m:32/n:32/k:32                6.18 us         6.03 us           32 2.03739G/s 10.8661G/s
    gemm/d/armpl/m:32/n:32/k:32                6.66 us         6.66 us           32 3.69217G/s 9.84578G/s
    gemm/s/armpl/m:128/n:128/k:128             33.6 us         33.7 us           32 5.84165G/s 124.622G/s
    gemm/d/armpl/m:128/n:128/k:128             34.1 us         33.9 us           32 11.6079G/s 123.817G/s
    gemm/s/armpl/m:512/n:512/k:512              579 us          579 us           32  5.4345G/s 463.744G/s
    gemm/d/armpl/m:512/n:512/k:512              892 us          892 us           32 7.05221G/s 300.894G/s

    Accelerate v1.11

    Benchmark                                     Time             CPU   Iterations       Byte       FLOP
    -----------------------------------------------------------------------------------------------------
    gemm/s/accelerate/m:32/n:32/k:32           4.76 us         4.84 us           32 2.53688G/s   13.53G/s
    gemm/d/accelerate/m:32/n:32/k:32           5.24 us         5.31 us           32 4.62607G/s 12.3362G/s
    gemm/s/accelerate/m:128/n:128/k:128        13.7 us         13.7 us           32 14.3641G/s 306.433G/s
    gemm/d/accelerate/m:128/n:128/k:128        28.7 us         28.9 us           32 13.6179G/s 145.257G/s
    gemm/s/accelerate/m:512/n:512/k:512         249 us          249 us           32 12.6255G/s 1.07738T/s
    gemm/d/accelerate/m:512/n:512/k:512         853 us          853 us           32 7.37947G/s 314.857G/s
  8. martin-frbg commented on Aug 11, 2026

    @martin-frbg
    Collaborator

    Thanks for the benchmark updates. In my latest tests, vlovero's kernels outclass the old NEON ones in single-threaded runs, it is only in multithreaded where the advantage of their single-threaded fast path over the classic blocked-GEMM algorithm is less pronounced (but also thermal throttling of my M4mini and the known issues with buffer assignment bottlenecks under load complicate the picture).
    I have put what I currently have into PR #5971 - @vlovero not sure if getting your code half-blindly plugged into OpenBLAS even was what you intended, but I cannot compete with you in this field, and you're credited in the Contributors list. (Also feel free to copy the PR and resubmit so that you're properly in the git history, if you agree with this at all)

  9. vlovero commented on Aug 12, 2026

    @vlovero
    Author

    @martin-frbg Everything seems good to me. Before a PR is merged I need to figure out a compiler dependent bug I'm experiencing. AppleClang seems to be optimizing away code blocks that handle the zero padding only when the size of the matrices are larger than the size of the static packing matrices. I've recompiled it with gcc and do not have the same issues. The sanitizers aren't catching anything and no compiler warnings so I'm not sure if it's UB or an actual compiler bug.

  10. vlovero commented on Aug 12, 2026

    @vlovero
    Author

    @martin-frbg I was able to get all of the tests to work on AppleClang with aggressive compiler optimizations enabled by using much more aggressive register clobbering. I'm not sure which register was the culprit (there's so many), but I think one of them may have been leading to some type of UB I don't quite understand.

    The updated code is on my repo now. I can duplicate your PR or do whatever is easiest if you think things are ready for OpenBLAS

  11. martin-frbg commented on Aug 13, 2026

    @martin-frbg
    Collaborator

    Thank you - I'll just add your changes to the clobber lists as I revamp the PR to include support for DYNAMIC_ARCH builds.
    (I've just tried with SGEMM, and indeed the need for the additional "clobber list only" inline assembly routines is resolved with your change.)
    One very minor issue that still bothers me is that the Fortran test codes (test/?blat3) complain that all of IEEE_INVALID_FLAG IEEE_DIVIDE_BY_ZERO IEEE_OVERFLOW_FLAG IEEE_UNDERFLOW_FLAG are signalling, suggesting that the fpsr got trashed by SME.
    Sneaking in an "msr fpsr, xzr\n\t" after the smstop takes care of it, but I'm not sure if that is a sane way to address it ? (I'm fairly certain that these condition flags are spurious, of course, but I wouldn't want to suppress them at the gfortran option level)

  12. vlovero commented on Aug 13, 2026

    @vlovero
    Author

    I think a manual clearing of the flags may be necessary; looking at ARM documentation says:

    When the Effective value of PSTATE.SM is changed by any method from 0 to 1 or from 1 to 0, the FPSR is set to the value 0x0000_0000_0800_009f, setting all the cumulative status bits to 1.

    And when I manually step through the assembly using lldb, every time an SME instruction gets called, the fpsr register gets changed back to 0x800009f (even if I set it to zero after entering streaming mode).

    I'll look into this more to see if there is a way to detect IEEE exceptions when using streaming.

  13. vlovero commented on Aug 13, 2026

    @vlovero
    Author

    Going through the ARM documentation it seems like it is not possible using the fpsr register because they are purposely not updated to maximize throughput

    RTGSKG: Produce the expected IEEE 754 default result but do not update the FPSR cumulative exception flag bits.

    So zeroing it out seems like the way to go to suppress the fortran test code complaints.

  14. vlovero commented on Aug 13, 2026

    @vlovero
    Author

    Performance was lacking compared to Accelerate and ArmPL for complex arithmetic. I changed to packing dimensions and got 8-20% speedups.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions