Skip to content

Add RISCV64_ZVL1024B (wide VLEN / SpacemiT K3 A100) - #6066

Draft
hugomeiland wants to merge 2 commits into
OpenMathLib:developfrom
hugomeiland:riscv64-zvl1024b
Draft

hugomeiland wants to merge 2 commits into
OpenMathLib:developfrom
hugomeiland:riscv64-zvl1024b

Conversation

@hugomeiland

@hugomeiland hugomeiland commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Draft PR to land a new RISCV64_ZVL1024B OpenBLAS target for RVV 1.0 cores with VLEN ≥ 1024. Developed and tested on the SpacemiT K3 (Banana Pi BPI-SM10), which has two asymmetric clusters:

Cluster Cores VLEN Typical OpenBLAS path
X100 0–7 256 (vlenb=32) existing RISCV64_ZVL256B — primary compute path on K3
A100 8–15 1024 (vlenb=128) new RISCV64_ZVL1024B — secondary / experimental on K3

Please keep this as a draft. On K3, A100 RVV is not the primary HPL/BLAS path (X100 + ZVL256B remains the workhorse; A100’s value is IME2 matrix accel, not plain RVV). This target is still useful for A100-only jobs and for heterogeneous X100+A100 MPI where A100 ranks need a VLEN-matched GEMM.

What this adds

  • New target + KERNEL.RISCV64_ZVL1024B, GEMM/TRMM from generate_kernel.py
    • Default tiles: dgemm/dtrmm 32×8, sgemm/strmm 64×8 (bitsliced A-pack), cgemm 8×8, zgemm 8×4
    • Power-of-two packers (gemm_*copy_{32,64}_rvv, TRMM/SYMM) with panels 32/16/…/1 matching bitsliced tails (odd-M safe)
    • TRSM MR=32/64 generic copies + SHIFT wiring in trsm_kernel_{LN,LT,RN,RT}.c
    • Optional contiguous A-pack helpers (gemm_*copy_contig_rvv.c) via generate_kernel.py a_pack=contiguous
    • Prior 16×8 kernels retained in-tree for A/B
  • -march=…_zvl1024b, DYNAMIC_ARCH dispatch: vlenb ≥ 128 → ZVL1024B, ≥ 32 → ZVL256B, else ZVL128B
  • CMake + README / docs/install.md wiring
  • Companion fix: stock ZVL256B dgemm/sgemm wide-vle+vget LMUL splits made VLEN-portable (assumed VLMAX@256 before), so ZVL256B kernels do not corrupt tails if run at larger VLEN

Important caveats (X100 vs A100)

  • Do not run a static TARGET=RISCV64_ZVL1024B library on X100 — kernels assume VLEN≥1024 and will misbehave / fault on VLEN=256.
  • Prefer DYNAMIC_ARCH=1 (one .so for both clusters) or per-rank FlexiBLAS backends.
  • On SpacemiT K3, A100 affinity requires registering the thread via /proc/set_ai_thread before exec; migrating after ld.so/OpenBLAS has cached VLEN decisions can SIGSEGV.
  • L1 and some helpers still share the ZVL256B/#else paths; SH/SB GEMM stay on the smaller tile for now.

Test plan

  • Static TARGET=RISCV64_ZVL1024B build on K3 (EESSI GCC 14 / foss/2025b)
  • Official OpenBLAS L2/L3 BLATs (S/D/C/Z blat2 + blat3) on A100 — all PASSED (incl. odd-M with PoT packers)
  • HPL (EESSI HPL 2.3, NB=192, FlexiBLAS) on BPI-SM10 @ N=12000:
    • X100×8 ZVL256B ≈ 52–53 GF (primary path)
    • A100×8 portable ZVL256B ≈ 15 GF
    • A100×8 ZVL1024B 32×8 ≈ 40.9 GF (~2.7× vs portable ZVL256B on A100; ~1.19× vs earlier 16×8 ≈35.9 GF)
    • Hetero 8×X100 (ZVL256B) + 8×A100 (ZVL1024B 32×8) ≈ 58.5 GF (beats X100-only; was ~57.5 with 16×8)
  • Upstream CI / other RVV boards: confirm DYNAMIC_ARCH still selects ZVL256B/ZVL128B correctly when vlenb < 128
  • Optional: CMake configure + build with -DTARGET=RISCV64_ZVL1024B

Board notes: https://www.opensolvers.com/boards/SM10.html
Patch provenance: https://github.com/opensolvers/benchmarks (OpenBLAS/)

Made with Cursor

hugomeiland and others added 2 commits September 27, 2026 15:59
Introduce GEMM/TRMM kernels generated for VLEN>=1024 (d/s 16x8, c 8x8,
z 8x4), DYNAMIC_ARCH dispatch on vlenb>=128, and CMake/docs wiring.
Also make stock ZVL256B GEMM vget splits VLEN-portable so the same
DYNAMIC_ARCH build remains correct on wider vector units.

Co-authored-by: Cursor <cursoragent@cursor.com>
Use dgemm 32x8 / sgemm 64x8 with power-of-two packers matching
bitsliced A-panel tails; add TRSM MR=32/64 copies and SHIFT wiring.
Default param MR stays 32/64; 16x8 remains available for A/B.

Co-authored-by: Cursor <cursoragent@cursor.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant