Repository navigation
apps/ggml: GGML quantized kernels on Approximations - #9511
Draft
alexreinking wants to merge 47 commits into
Draft
alexreinking wants to merge 47 commits into
alexreinking wants to merge 47 commits into
Conversation
alexreinking
added this pull request to stack #9512
October 6, 2026 17:47
alexreinking
removed this pull request from stack #9512
October 6, 2026 19:33
alexreinking
force-pushed
the
alexreinking/ggml-quant
branch
from
October 6, 2026 19:34
0542a85 to
3f8e310
Compare
alexreinking
added this pull request to stack #9522
October 6, 2026 19:34
alexreinking
force-pushed
the
alexreinking/ggml-quant
branch
2 times, most recently
from
October 6, 2026 20:16
1627d58 to
9ed2641
Compare
alexreinking
force-pushed
the
alexreinking/ggml-quant
branch
from
October 7, 2026 06:39
9ed2641 to
401c635
Compare
alexreinking
removed this pull request from stack #9522
October 7, 2026 16:08
alexreinking
added this pull request to stack #9524
October 7, 2026 16:08
alexreinking
force-pushed
the
alexreinking/ggml-quant
branch
from
October 7, 2026 16:08
097c0b7 to
cb9cd78
Compare
alexreinking
force-pushed
the
alexreinking/ggml-quant
branch
from
October 7, 2026 17:26
cb9cd78 to
68152e9
Compare
alexreinking
force-pushed
the
alexreinking/ggml-quant
branch
from
October 7, 2026 17:47
68152e9 to
f631a55
Compare
alexreinking
removed this pull request from stack #9524
October 7, 2026 19:21
alexreinking
force-pushed
the
alexreinking/ggml-quant
branch
from
October 7, 2026 19:26
f631a55 to
74a1b3f
Compare
alexreinking
added this pull request to stack #9527
October 7, 2026 19:26
alexreinking
removed this pull request from stack #9527
October 8, 2026 11:09
alexreinking
force-pushed
the
alexreinking/ggml-quant
branch
from
October 8, 2026 11:10
74a1b3f to
7bc3997
Compare
alexreinking
added this pull request to stack #9529
October 8, 2026 11:10
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
blocked_mmla now tiles nt x mt outputs as 2 x 2 smmla sub-tiles and splits the storage of the rfactor intermediates (acc, blk, codes) by 2 in n and m, so each sub-tile is 4 dense lanes and the accumulators stay in registers (row-major storage spilled: 112 stack refs per 4 x 8 block). The algorithm is unchanged. Inner loop (q4_0 x q8_0 and f32-to-q8_0): 197 instrs per 4 x 8 x 32 block, 32 smmla, no stack traffic (2 x 8: 137 per 16 smmla). Needs the split_storage vectorized-index fix (ad2c2ed). vec_dot, codec and float mul_mat asm unchanged. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
out(n, m) = 0; out += ... zeroed the whole output serially (a _bzero before the parallel loop) and read-modified-wrote out per tile; this was the 1-4% gap between split_storage (d) and the blocked-coordinates prototype (b), whose inner loops are identical up to register allocation. Now dot(n, m) += W * X and out(n, m) = dot(n, m), with dot computed per output tile. No zeroing pass; the mmla inner loop is 191 instrs / 32 smmla / 0 stack refs (was 197). A pure out also allows ShiftInwards on its tiles, so the gemm specializations require N >= nt, M >= mt instead of divisibility: any token count M >= 8 runs full 4 x 8 smmla tiles (edge tiles overlap) instead of falling back to 4 x 4 sdot or gemv pairs. Sub-tile splits in blocked_mmla are GuardWithIf: its intermediates also appear (never run) in the other tiles' branches, whose bounds must not grow. vec_dot asm unchanged except q4_0_f32_to_q8_0 (no load/add of the zeroed output; fewer spills). --check: 0 failures, incl. odd and M-tail shapes. Timing pending AC. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
apps/ggml/runtime/thread_pool_common.h is src/runtime/thread_pool_common.h copied verbatim from upstream main at ba4b018, its last change there (blob f8ffaa1, the same as on this branch's base), so that the app can try thread-pool changes without changing Halide's runtime; diffing the two files shows every change. runtime/thread_pool.cpp compiles it as ordinary C++. The runtime defines all thread-pool entry points weakly, so the copy defines them strongly (halide_do_par_for, halide_do_parallel_tasks, halide_semaphore_*, halide_set/get_num_threads, halide_shutdown_thread_pool, the halide_default_* handlers, ...), and the linker binds both the kernels' and the runtime's own calls to it. Unlike halide_set_custom_parallel_runtime, this also covers the thread count and shutdown, and needs no registration at startup. The glue provides what runtime_internal.h and runtime_atomics.h give runtime code, and wraps the pool's C++ internals in a namespace of its own, so they can't collide with the runtime's weak copies. It is an OBJECT library: from an archive it would never be loaded, since the runtime's weak definitions already satisfy every reference. No behavior change. A link map of ggml-quant-bench resolves halide_do_par_for, halide_do_parallel_tasks, halide_semaphore_release, halide_set_num_threads and halide_shutdown_thread_pool to thread_pool.cpp.o; --check: 454 results, 0 failures. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Applies the thread_pool_common.h part of the upstream keep-awake proposal (#9526, e9ac133) verbatim to the app's copy of the pool: a refcounted count; while it is held, up to num_threads - 1 idle workers and owners waiting on their own loops poll for work instead of sleeping, idle workers are not demoted to the B team, an idle thread parks anyway after 4096 polls without work, and acquiring the first reference wakes both teams. The count survives halide_shutdown_thread_pool. The glue renames halide_thread_pool_keep_awake to ggml_halide_thread_pool_keep_awake, so it can't collide with the runtime's if #9526 lands, and runtime/thread_pool.h declares it with an RAII holder, ggml_halide::ThreadPoolKeepAwake. ctest ggml_thread_pool (runtime/thread_pool_test.cpp, after #9526's thread_pool_keep_awake_aottest) stress-tests the copy through the q4_0 x q8_0 mul_mat kernel (exact integer results) and direct calls: nested loops, do_parallel_tasks with semaphores, error propagation, keep-awake on and off with gaps, thread-count changes, shutdown/restart, and concurrent callers. No timing assertions. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
With keep-awake, what remains of fork/join cost is the work queue's mutex, taken to queue a loop, to claim and finish each iteration, and to wait. A top-level halide_do_par_for started on an idle pool (nothing queued, enough idle A-team workers polling) now publishes the loop in a global slot instead. Polling workers join it and claim iterations from an atomic counter. The owner runs iterations too, closes the loop, and then waits for its helpers to leave (an active count, checked Dekker-style against the loop's state). Helpers are limited to min(size, threads) - 1, the first error is returned, and loops with one iteration or one thread run inline. Everything else takes the existing path: a loop started while the slot is in use (nested in a fast-path iteration, or from another thread), a busy pool, and halide_do_parallel_tasks (semaphores, async). The glue gains the seq_cst atomics the change uses, plus ggml_halide_thread_pool_fast_loops() for tests and diagnostics. The test gains fast-path cases. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
GGML's threadpool is persistent: it is resumed before each sample and paused after it, and within a sample its workers spin at barriers and poll for the next graph. To match, HalideKernel::run holds the app pool's keep-awake count for the duration of each CPU sample, so idle Halide workers poll for the next parallel loop rather than sleeping, and they are let go between samples. --no-keep-awake turns this off for ablation. The README's methodology section describes both sides. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The f32 -> q8_0 activation encode (compute_root before the dot) was scalar: a scalar max loop and 32 scalar code stores per block. The new generic `encoder` schedule computes each encode-side intermediate per record, vectorizing reductions (atomic, across the block: vector_reduce_max) and per-value Funcs across their values. It needs no per-format lines. round_away (C's roundf) is now trunc(v + copysign(0.49999997f, v)), equal to roundf for every float (checked exhaustively), instead of trunc and a compare/select. The q8_0 encode loop goes from scalar to 79 NEON instrs per block (111 with the old round_away). vec_dot (q8_0 acts) asm is unchanged; --check 454/0, q4_0 mul_mat full 384/0, odd shapes 216/0. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
<type>_quantize computes each encoded block per output block and schedules it with mul_mat's `encoder` (no per-format lines). `encoder` now vectorizes every stage of a per-value Func across its values (q4_0's planar_field_bytes OR-packs nibbles in an update), and atomic-vectorizes only per-record reductions (q8_0's maximum; q4_0's Tuple argmax stays a serial loop, as GGML's scalar q4_0 reference). Single-threaded, AC, median of 3 interleaved reps (ns; GGML from_float): q8_0 4096 1772 -> 757 (677); q8_0 14336 6167 -> 2659 (2344); q4_0 4096 3857 -> 2376 (2156); q4_0 14336 13548 -> 8439 (8037). The rest of q8_0's gap is round_away (3 extra instrs per 4 lanes; an fcvtas peephole in core would remove them). Only the two quantize libraries' asm changes; --check 454/0, ctest 2/2. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
A thread that owns a job may run a serial task belonging to another parallel region. When that task's semaphore runs dry, the thread puts it back and, if its own job is finished, returns without looking at it again. Meanwhile the task's owner may have seen the task in use and gone to sleep: a semaphore release wakes it, but it still can't add itself to a serial task that is in use. Nobody then runs the rest of the task. Putting a serial task back now wakes its sleeping owner. The bug is upstream's (ba4b018), but the fast path makes it likely: in ggml_thread_pool_test, the fast-path owner and its helpers each run halide_do_parallel_tasks at the same time, and the test hung about once in 30 runs with HL_NUM_THREADS=5. With the fix, 100 runs pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
On the fast path, every participant claimed each iteration from one shared counter, counted itself in and out of a shared active count, and the owner waited for every helper to leave before returning. At T >= 8, workers spent much of their time on those contended lines. Each participant (the owner, and each worker, numbered as it starts) now has its own cache line, with a presence flag and a reserved first iteration. It claims that, then claims from the shared counter (which starts past the reserved iterations, and is read before it's incremented), then takes over reserved iterations whose participants haven't shown up, so a late helper never holds up the loop. It adds what it claimed to a shared total once, and the owner returns when the total reaches the loop's size. Waiting for stale helpers to leave moves to the start of the next fast-path loop, by when they're usually long gone. Based on critical-path-2's runtime_ticket.patch prototype. The ctest gains a test of slow and late helpers, oversubscription and HL_NUM_THREADS from 1 to 64. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Claiming from one shared counter still made every participant write the same cache line for every iteration, and handed iterations out in whatever order threads arrived, whereas GGML gives each thread a fixed range of rows. The fast path now splits a loop into min(size, threads) contiguous chunks, one per participant: the owner's, then one per worker in the order the workers started. Each chunk's counter sits on its participant's cache line, next to its presence flag. A participant claims from its own chunk, then from the others' in turn, so a late or absent helper's iterations are run by whoever is free: unlike runtime_affinity.patch, where only the owner took them over, serially, after closing the loop, which stalled at T=12. A worker also gets the same iterations, and so the same data, from one loop of a given size to the next. The ctest adds lowering the thread count below the number of workers, so that some workers that join a loop have no chunk of their own. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
With 32-row tasks, N = 512 gives 16 tasks: two rounds on 10-12 threads, so gemv stopped scaling past T = 8. 16-row tasks (32 at N = 512) balance 8-12 threads now that the app's pool claims are cheap. AC, 3 interleaved reps, T 8/10/12, best-T geomean vs ggml-cpu over 2048x512, 3584x512, 4096x1024, 2048x2048 (32-row -> 16-row): hot q8_0 0.938 -> 1.040, hot f32-(b) 0.906 -> 0.949, cold q8_0 0.999 -> 0.993, cold f32-(b) 0.994 -> 0.991. T-sized tasks (rows = ceil(N / (k * halide_get_num_threads()))) were no better: hot q8_0 1.002 (k = 1), 0.978 (k = 2); hot f32 0.946, 0.920. Only the mul_mat libraries' task split changes; --check 454/0, q4_0 mul_mat full 384/0, odd shapes 216/0, ctest 2/2. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Replace the app pool's lost-wakeup condition with the one from #9534. The earlier fix only woke the put-back task's own owner, but that owner may be asleep on a sibling task, so the group could still hang. Wake all sleeping owners instead, as a semaphore release does. Add #9534's regression test to ggml_thread_pool_test. It forces a thread to steal a serial task while its own work finishes. It uses a parallel task rather than halide_do_par_for, which would usually take the fast path. It fails 10 of 10 runs with the old condition and passes with the new one. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
A loop started right after another (e.g. a parallel activation encode right before the gemv) cost about 1-1.7 us more than the first at T=8-12, because the owner reset shared counters, waited for every helper to leave, and helpers yielded between polls. - Tag each chunk's claim and done counters with the loop, so a new loop needs no reset and the owner returns when every chunk is done, without waiting for helpers to leave; a stale helper can't claim from a newer loop. - Read the slot's description like a seqlock; close the previous loop lazily when the next one opens. - While keep-awake is held, workers that took part in a loop poll with the pause instruction (no yield) for 2^15 polls. - Keep awake only workers numbered below threads - 1, so after the thread count drops the polling workers are the ones with chunks. Extend ggml_thread_pool with sequences of back-to-back loops (extents from 1 up, HL_NUM_THREADS 2/5/default/64, keep-awake held or not, failing, sleeping and nested iterations, thread-count changes). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… repack
A format type may now name two lossless stages composed after the
quantizer, chosen independently in formats.cmake:
<type>[.<codes>][.<rows>x<chunk>]. Codes i4 are two's-complement nibbles
(twos, schemes/common.h). The layout Interleave{rows, chunk}
(schemes/layouts.h, domain-free) stores one record per `rows` rows, with
each field's elements interleaved across the rows in `chunk`-byte pieces.
q4_0.i4.4x8 is GGML's repacked q4_0_4x8.
The codec and mul_mat generators are layout-polymorphic: weights are
[K / block, N / rows] records, and kernels/ has no per-layout code. The
harness repacks the weights with a test-only transcription of GGML's
repack (harness/repack.cpp, MIT, attributed). --check compares the codec
bitwise against that transcription for 4x4, 4x8 and 8x8, and compares the
transcription against the bytes in GGML's own CPU_REPACK buffer for the
layout it picks here (4x8). The AoS libraries' asm is unchanged.
The interleaved mul_mat is correct but not yet scheduled for its layout
(no smmla yet).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Bound out's N by w's record count times the rows per record, instead of deriving w's extent from out's. N is then a known multiple of the interleave, so shifted-inwards tile origins are visibly aligned and the interleaved 4x8 gemm's loads stay dense (32 smmla, no spills, no gathers). AoS libraries are byte-identical. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
An act whose compute format has a layout (f32:q8_0.4x8) splits M: out's first update takes the rows in whole records from dotr over X in that format, on gemm tiles starting on a record (guarded M tails; dotr's rows past the records read the last row, clamped), its second the remaining rows from dot over X in its AoS form, on gemv tiles. Registered for q4_0.i4.4x8. The clamp-only alternative (one sum over the interleaved act) was 7x slower at M = 1: its gemv gathers from the interleaved act. The comparison vs GGML's repack/KleidiAI is pending a quiet machine. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
A laid-out integer weight's gemv now runs one record per tile, as GGML's repacked gemv: int32 sums vectorized across the record's rows and each piece's 4-code quads (sdot), so each piece is one dense load and the activation's piece is broadcast. The steady-state loop is 86 instructions per 2 blocks (16 sdot, no spills, no gathers). Tiles of whole records use RoundUp (exact: N is a whole number of records), since ShiftInwards hides the tile base's alignment behind likely_if_innermost (gap L). The preserved RVar is named ryi_rows: dot's specializations share loop bounds by RVar name (gap M, a core bug reported separately with a GGML-free repro). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
With interleaved activations of at least 16 rows, gemm runs 4 x 16 smmla tiles whose codes and scaled blocks are computed per row pair, so LLVM shares the weight's decode across the tile (203 instructions per 64 outputs, vs 234 on 4 x 8). Fewer rows keep the 4 x 8 tile; dot's bound is a select on the specializations' simple conditions, so each branch's allocation stays constant. AoS (non-record) tiles are unchanged. T = 8, AC, vs GGML's fastest in the same invocation, gate 2b f32 gemm rows (M = 32/512): median 0.89 -> 0.94, min 0.84 -> 0.88. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Q4_0Quant gains a power-of-two code_shift (default 0): encode emits
(c << k, s / 2^k), decode is c'·s'. The q4_0 factory's codec slot holds
Parallel{codes: twos(4)} by default and Parallel{codes: twos16,
scale: Scale(16)} for the shifted variant; the packers and the layout
are unchanged, so both encode to identical bytes. Exact barring
subnormal f32 scales (documented on Q4_0Quant).
The codec check builds the scaled quantize/dequantize per codec row and
compares bytes/values with the reference, plus a crafted full-range
dequantize (all fp16 scales, all code bytes) vs GGML's to_float.
Kernels still use the reference; the per-kernel choice is pending AC
timing.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
`q4_0.soa` keeps the planar packer's ports apart (layout "soa" is Identity): codes [16, K/32, N] and fp16 scales [K/32, N] are separate kernel arguments. matmul.cpp and the codec take one input per encoded port (layout-polymorphic; AoS asm byte-identical, 15/15). The harness calls mul_mat and codecs through _argv, builds port buffers from the metadata, and checks SoA against GGML's blocks split alike (relayout). --check 554/0 (full 864/0, odd 420/0); ctest 2/2; kernels/ 433 lines. SoA vs AoS timing pending AC. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… tiles, gemm tasks per thread Picks from AC timings (same-invocation geomeans, T = 8/10/12): - mul_mat on multi-row weight records (q4_0.i4.4x8) takes A' (codes at 2^4, the exact Q4_0Quant code_shift alternative); its 4 x 16 smmla tiles then compute the codes per block (at mio they spill): gemv 1.06/1.14/1.10x, gemm 1.04/1.03/1.06x vs the reference decode. AoS libraries keep the reference (A' measured 0.95-0.97x on their gemv). - 4 x 4 smmla tiles where the 4 x 16 / 4 x 8 ones would waste 4 or 12 rows, below 64 rows (Mr = 4, 12, 20, 36, 52): odd-M geomean 1.18-1.26x; Mr = 28 and 100 measured slower and keep the taller tiles. - gemm tasks shrink to leave at least 2 per thread (runtime thread count): 1.005-1.13x across T = 8/10/12 on every library. asm (interleaved lib): gemv 78 instrs / 16 sdot; 4 x 16 220 / 64 smmla / 12 spills; 4 x 8 115 / 32 / 0; 4 x 4 68 / 16 / 0. vec_dot and codec asm unchanged. --check 554/0 (q4_0 full 864/0, odd/edge 420/0), ctest 2/2. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Adds mul_mat to the q4_0.i4.4x4 row (GGML's block_q4_0x4: 4 rows in 4-code pieces). blocked_rows now keeps the act's integer codes out of the eager inline and stages them per block in a vectorized register Func (a.in(quads).compute_at(quads, u)), so each 4-code act piece is a lane of a register and LLVM emits by-element sdot (sdot vD.4s, vW.16b, vA.4b[i]) instead of an ld1r.4s broadcast per piece. gemv steady state (2 blocks x 4 rows), instrs / sdot / spills: 4x4 (f32:q8_0.4x4): 90 -> 62 / 16 by-element / 0 (C++ spike: 55-57) 4x8 (f32:q8_0.4x8): 78 -> 74 / 16 / 0 vec_dot, codec and AoS mul_mat asm byte-identical. --check 572/0; full q4_0 1152/0; odd/edge 480/0; extra 4x4/4x8 shapes 108/0; ctest 2/2. Timing vs KleidiAI / repack pending AC. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…gemm The schedules take approx[0] as the weight's; configure() now throws if the weight has no scheme, the only way that order could break. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The interleaved 4x4 layout's gemm now follows KleidiAI's dotprod gemm: per block the weight record is decoded once, and int32 sums are computed per activation record by by-element sdot. The record's codes are staged in memory order, so one 16-byte load holds a piece of each of its 4 rows, and its 4 scales are staged as one vector. On AC (M3, T = 8/10/12, 7 gemm rows, 3 rounds, same invocation as GGML) the 4x4 lib's gemm is 1.16-1.19x faster (geomean) than before. It is at parity with the 4x8 smmla gemm (paired 0.99/1.00/1.02), so 4x8 stays the gemm default. The 4x8 and AoS asm are unchanged. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
A weight is repacked once (as KleidiAI's is), so q4_0.i4.4x4 is the default for gemv and gemm; 4x8 (smmla) and AoS stay comparison libraries. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
mul_mat's reduction is now RDom(0, block, 0, blocks): r.x a value within a block, r.y the block. The operands are read through a block view F(j, b, i) = flat(b * block + j, i), and formats in the weight's blocks approximate that view with a block-indexed BlockReshape, so no index into a block is a % or / of the reduction (the schedules use r.x and r.y instead of splitting a flat r). The encoded ports, ABI and bytes are unchanged; codec.cpp keeps the flat formats. Prerequisite for staging the act per 4-code piece (gap N). Asm: all libraries identical or relabelled except the 4x4 f32:q8_0.4x4 gemm, whose 16-row loop is 261 instructions / 18 spill+reload (was 266 / 15); no timing yet. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
alexreinking
force-pushed
the
alexreinking/ggml-quant
branch
from
October 10, 2026 13:22
8edc69e to
5ab903f
Compare
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Work in progress; top of stack #9540. Reimplements GGML's quantized
vec_dotandmul_matin Halide onFunc::approximate_by()(#9486), and compares them against GGML's CPU paths (plain, repack, KleidiAI) on Neon. Metal comes later. Touches onlyapps/. Replaces #9417.ggml-quant-bench): GGML v0.26.0 through a vcpkg port, reached only through public headers. An f64 oracle with tolerances derived from each provider's declared precision; randomized provider order; medians with 95% CIs;--coldfor weights streamed from DRAM.mul_matalgorithm,out(n, m) = sum_k W(k, n) * X(k, m), serves q8_0, f16 and f32 activations, withvec_dot, gemv and gemm as schedules. Weight layouts: GGML's plain records, its 4x4 and 4x8 repacked (AoSoA) records, and an SoA layout, all expressed as Approximations. Onarm_i8mmtargets the gemm usessmmlaon 4x16, 4x8 or 4x4 register tiles whose accumulatorssplit_storagekeeps in registers; the gemv on 4x4 records uses by-elementsdot(Test by-element sdot/udot codegen on ARM #9544).kernels/is 477 lines, with no per-format code.Results so far (M3 Max, AC power).
vec_dot, 1 thread: GGML/Halide is 0.93 at n = 256 (adapter call overhead) and 0.99-1.00 at n = 4096 and 14336.mul_matwith q4_0 weights on 60 Llama 3, Llama 3.2 and Qwen 2.5 layer shapes, warm and cold, each side at its best thread count (GGML/Halide, >1 means Halide is faster; ranges, median in parentheses):Against plain GGML (the same weight layout), every row is faster or within the 95% CIs. The app links its own copy of Halide's thread pool (
apps/ggml/runtime/), whose lock-freedo_par_forfast path and keep-awake bring fork/join at 8-12 threads down from 22-40 us to 1-1.5 us.With interleaved weights (preliminary; both AC runs were noisy, load 2-17): against GGML's fastest f32-activation path, all 24 gemv rows are at or above 1.0 (geomean 1.17); gemm geomean is 1.08 at M = 8 and 0.99 otherwise, with 7 rows at 0.93-0.97 (M = 8 and 32, against repack or KleidiAI).
Tests: ctest
ggml_check(ggml-quant-bench --check), every GGML type against the oracle plus the Halide codecs and kernels, 0 failures;ggml_thread_pool, the vendored pool.Authored by GitHub Copilot