-
Notifications
You must be signed in to change notification settings - Fork 21.1k
Pull requests: ggml-org/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
quant : do not require imatrix when the tensor keeps its type
#26255
opened Jul 28, 2026 by
TrevorS
Contributor
Loading…
chat : add qwen3 specialized parser
testing
Everything test related
#26252
opened Jul 28, 2026 by
aldehir
Contributor
Loading…
Add DMMV ESIMD Q3_K kernel
documentation
Improvements or additions to documentation
ggml
changes relating to the ggml tensor library for machine learning
SYCL
https://en.wikipedia.org/wiki/SYCL - GPU programming language
Refactor/extract model resolution
testing
Everything test related
#26247
opened Jul 28, 2026 by
ServeurpersoCom
Contributor
Loading…
ggml: fix integer overflow in tensor size computation (ggml_new_tensor_impl)
ggml
changes relating to the ggml tensor library for machine learning
#26245
opened Jul 28, 2026 by
nskath
Loading…
vendor: update BoringSSL to 0.20260728.0
vendor
#26241
opened Jul 28, 2026 by
cabelo
Contributor
Loading…
mtmd: bound InternVL preproc_max_tiles read from GGUF
mtmd
Related to multimodal functionality (video/image/audio)
#26237
opened Jul 28, 2026 by
mtholmquist
Loading…
openvino: phase-split prefill/decode with USM host KV (CPU + iGPU)
documentation
Improvements or additions to documentation
ggml
changes relating to the ggml tensor library for machine learning
OpenVINO
testing
Everything test related
#26235
opened Jul 28, 2026 by
lslusarczyk
Contributor
•
Draft
3 of 4 tasks
[SYCL] support dev2dev memcpy by host forward
documentation
Improvements or additions to documentation
ggml
changes relating to the ggml tensor library for machine learning
SYCL
https://en.wikipedia.org/wiki/SYCL - GPU programming language
#26234
opened Jul 28, 2026 by
arthw
Contributor
Loading…
model: Align Laguna-S-2.1 chat template to huggingface
testing
Everything test related
#26232
opened Jul 28, 2026 by
crusaderky
Contributor
Loading…
[SYCL] Support q2 mul_mat
ggml
changes relating to the ggml tensor library for machine learning
SYCL
https://en.wikipedia.org/wiki/SYCL - GPU programming language
#26231
opened Jul 28, 2026 by
arthw
Contributor
Loading…
mimo2: add MTP draft support
conversion
model
Model specific
#26228
opened Jul 28, 2026 by
tnhnyzc
Contributor
Loading…
core : support output vocab size distinct from embedding vocab size
conversion
#26226
opened Jul 28, 2026 by
adithyab94
Contributor
Loading…
2 tasks done
Proper fix for host buffer sync
ggml
changes relating to the ggml tensor library for machine learning
#26225
opened Jul 28, 2026 by
pwilkin
Member
Loading…
metal: fix NaN in mul_mm_id when activations exceed f16 range
Apple Metal
https://en.wikipedia.org/wiki/Metal_(API)
ggml
changes relating to the ggml tensor library for machine learning
testing
Everything test related
#26223
opened Jul 28, 2026 by
mdegans
Contributor
Loading…
convert : fix bytes_to_unicode import for transformers >= 5.15
conversion
#26217
opened Jul 28, 2026 by
SolshineCode
Loading…
server : add /slots endpoint action=clone_to (KV clone between slots)
documentation
Improvements or additions to documentation
server
#26204
opened Jul 27, 2026 by
solethais
Loading…
HIP: MMQ Dispatch config modification - separation of RDNA3, 3.5 from 4 and tune 4.
CUDA
Related to the CUDA backend
ggml
changes relating to the ggml tensor library for machine learning
#26199
opened Jul 27, 2026 by
Geramy
Loading…
server: fix prompt cache entry selection and f_keep filter
documentation
Improvements or additions to documentation
server
#26198
opened Jul 27, 2026 by
q-g-j
Loading…
ggml-cpu: add -mavxvnni for clang-cl when GGML_AVX_VNNI is enabled
ggml
changes relating to the ggml tensor library for machine learning
#26187
opened Jul 27, 2026 by
MaxCrazy1101
Loading…
model: add Kimi-K3 text model
conversion
model
Model specific
testing
Everything test related
#26185
opened Jul 27, 2026 by
pwilkin
Member
Loading…
feat: Add '--ignore-api-samplers' CLI flag to ignore API sampler settings
server
#26182
opened Jul 27, 2026 by
Sciguy429
Loading…
Support quantized kv cache for Minimax M3
model
Model specific
#26180
opened Jul 27, 2026 by
timkhronos
Contributor
Loading…
Add more benchmarks to llama-eval
documentation
Improvements or additions to documentation
examples
#26174
opened Jul 27, 2026 by
pwilkin
Member
Loading…
Previous Next
ProTip!
Find all pull requests that aren't related to any open issues with -linked:issue.