Add RISCV64_ZVL1024B (wide VLEN / SpacemiT K3 A100) - #6066
Draft
hugomeiland wants to merge 2 commits into
Draft
hugomeiland wants to merge 2 commits into
hugomeiland wants to merge 2 commits into
Conversation
Introduce GEMM/TRMM kernels generated for VLEN>=1024 (d/s 16x8, c 8x8, z 8x4), DYNAMIC_ARCH dispatch on vlenb>=128, and CMake/docs wiring. Also make stock ZVL256B GEMM vget splits VLEN-portable so the same DYNAMIC_ARCH build remains correct on wider vector units. Co-authored-by: Cursor <cursoragent@cursor.com>
Use dgemm 32x8 / sgemm 64x8 with power-of-two packers matching bitsliced A-panel tails; add TRSM MR=32/64 copies and SHIFT wiring. Default param MR stays 32/64; 16x8 remains available for A/B. Co-authored-by: Cursor <cursoragent@cursor.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Draft PR to land a new
RISCV64_ZVL1024BOpenBLAS target for RVV 1.0 cores with VLEN ≥ 1024. Developed and tested on the SpacemiT K3 (Banana Pi BPI-SM10), which has two asymmetric clusters:vlenb=32)RISCV64_ZVL256B— primary compute path on K3vlenb=128)RISCV64_ZVL1024B— secondary / experimental on K3Please keep this as a draft. On K3, A100 RVV is not the primary HPL/BLAS path (X100 + ZVL256B remains the workhorse; A100’s value is IME2 matrix accel, not plain RVV). This target is still useful for A100-only jobs and for heterogeneous X100+A100 MPI where A100 ranks need a VLEN-matched GEMM.
What this adds
KERNEL.RISCV64_ZVL1024B, GEMM/TRMM fromgenerate_kernel.pydgemm/dtrmm32×8,sgemm/strmm64×8 (bitsliced A-pack),cgemm8×8,zgemm8×4gemm_*copy_{32,64}_rvv, TRMM/SYMM) with panels 32/16/…/1 matching bitsliced tails (odd-M safe)SHIFTwiring intrsm_kernel_{LN,LT,RN,RT}.cgemm_*copy_contig_rvv.c) viagenerate_kernel.py a_pack=contiguous-march=…_zvl1024b,DYNAMIC_ARCHdispatch:vlenb ≥ 128→ ZVL1024B,≥ 32→ ZVL256B, else ZVL128Bdocs/install.mdwiringdgemm/sgemmwide-vle+vgetLMUL splits made VLEN-portable (assumed VLMAX@256 before), so ZVL256B kernels do not corrupt tails if run at larger VLENImportant caveats (X100 vs A100)
TARGET=RISCV64_ZVL1024Blibrary on X100 — kernels assume VLEN≥1024 and will misbehave / fault on VLEN=256.DYNAMIC_ARCH=1(one.sofor both clusters) or per-rank FlexiBLAS backends./proc/set_ai_threadbefore exec; migrating afterld.so/OpenBLAS has cached VLEN decisions can SIGSEGV.#elsepaths; SH/SB GEMM stay on the smaller tile for now.Test plan
TARGET=RISCV64_ZVL1024Bbuild on K3 (EESSI GCC 14 /foss/2025b)blat2+blat3) on A100 — all PASSED (incl. odd-M with PoT packers)vlenb < 128-DTARGET=RISCV64_ZVL1024BBoard notes: https://www.opensolvers.com/boards/SM10.html
Patch provenance: https://github.com/opensolvers/benchmarks (OpenBLAS/)
Made with Cursor