forked from ggml-org/llama.cpp
-
-
Notifications
You must be signed in to change notification settings - Fork 389
Pull requests: TheTom/llama-cpp-turboquant
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
fattn-vec: split the turbo K dot at D>128 to stop the VGPR spill (#294) — needs AMD validation
CUDA
ggml
#295
opened Aug 13, 2026 by
TheTom
Owner
Loading…
Fix MSVC C2159: duplicate extern from GGML_API on Windows
ggml
#292
opened Aug 12, 2026 by
sroller
Loading…
Pr/prefetch weights
AMD ZenDNN
Apple Metal
Ascend NPU
conversion
CUDA
devops
documentation
Improvements or additions to documentation
examples
ggml
Hexagon
IBM zDNN
jinja parser
model
mtmd
OpenCL
OpenVINO
server/ui
server
SYCL
testing
vendor
Vulkan
WebGPU
Pr/metal vulkan v1
AMD ZenDNN
Apple Metal
Ascend NPU
conversion
CUDA
devops
documentation
Improvements or additions to documentation
ggml
Hexagon
IBM zDNN
jinja parser
model
mtmd
OpenCL
OpenVINO
server/ui
server
SYCL
testing
vendor
Vulkan
WebGPU
#289
opened Aug 10, 2026 by
giveen
Loading…
Fix determinism leaks in TurboQuant KV cache and WHT state management
documentation
Improvements or additions to documentation
ggml
#281
opened Aug 8, 2026 by
giveen
Loading…
docs : add harness init strided-view bug to coverage limits
documentation
Improvements or additions to documentation
#278
opened Aug 8, 2026 by
giveen
Loading…
[ggml-cuda] Compute the turbo4 centroid LUT in advance during the decoding process
CUDA
ggml
#267
opened Aug 5, 2026 by
ghkang-cs
Loading…
hip: VEC flash-attn for D=512 (Gemma 4) on ROCm with quantized KV
ggml
Nvidia GPU
#156
opened May 24, 2026 by
cclecle
Loading…
vulkan: add TurboQuant KV cache support and optimized turbo mat-vec paths
ggml
Vulkan
#140
opened May 10, 2026 by
Fenix46
Loading…
fix(qwen35): support Qwen3.5:9B loading from Ollama GGUF
model
#135
opened May 8, 2026 by
Jordan-HS
Loading…
vendor: bump cpp-httplib to 0.43.2 (openssl 4.0.0 fix)
python
script
#121
opened May 4, 2026 by
TheTom
Owner
Loading…
1 of 3 tasks
HIP mixed TurboQuant vec FA on gfx900/gfx906
build
ggml
Nvidia GPU
#99
opened Apr 21, 2026 by
2bigO
Loading…
perf: turbo VEC flash attention — +9% decode on CUDA via autoresearch
ggml
Nvidia GPU
script
#53
opened Apr 4, 2026 by
signalnine
Loading…
7 tasks done
fix: HIP/ROCm compatibility — check cudaMemcpyToSymbol errors, guard …
ggml
Nvidia GPU
#41
opened Apr 1, 2026 by
terrysimons
•
Draft
ProTip!
Exclude everything labeled
bug with -label:bug.