llama.cpp

mirror of https://github.com/ggml-org/llama.cpp.git synced 2026-06-30 08:10:20 +00:00

Files

T

Aman Gupta 8bece2eb20 CUDA: use mmvq for mul-mat-id for small batch sizes (#18958 )

* CUDA: use mmvq for mul-mat-id for small batch sizes

* add mmvq too

* Fix perf issue on ampere. Use mmvf mm-id only for non-nvidia GPUs

* templatize multi_token_path

2026-02-03 23:31:23 +08:00

ggml-blas

ggml : add ggml_build_forward_select (#18550 )

2026-01-19 20:03:19 +02:00

ggml-cann

docs : Minor cleanups (#19252 )

2026-02-02 08:38:55 +02:00

ggml-cpu

ggml-cpu: FA split across kv for faster TG (#19209 )

2026-02-03 01:19:55 +08:00

ggml-cuda

CUDA: use mmvq for mul-mat-id for small batch sizes (#18958 )

2026-02-03 23:31:23 +08:00

ggml-hexagon

ggml-hexagon: flash-attention and reduce-sum optimizations (#19141 )

2026-01-30 21:14:20 -08:00

ggml-hip

HIP: add mmf for CDNA (#18896 )

2026-01-29 11:10:53 +01:00

ggml-metal

metal : minor cleanup (#19251 )

2026-02-03 13:43:29 +02:00

ggml-musa

CUDA: faster tile FA, add oob checks, more HSs (#16492 )

2025-10-11 20:54:32 +02:00

ggml-opencl

opencl: refactor some ops, concat, repeat, tanh and scale (#19226 )

2026-02-02 15:54:43 -08:00

ggml-rpc

rpc : use unordered_map::reserve and emplace (#18513 )

2026-01-02 12:09:36 +02:00

ggml-sycl

Remove support for Nvidia & AMD GPU, because the oneAPI plugin for Nvidia & AMD GPU is unavailable: download/installation channels are out of work. (#19246 )

2026-02-02 21:06:21 +08:00

ggml-virtgpu

ggml: new backend for Virglrenderer API Remoting acceleration (v2) (#18718 )

2026-01-28 17:49:40 +08:00

ggml-vulkan

docs : Minor cleanups (#19252 )

2026-02-02 08:38:55 +02:00

ggml-webgpu

Remove pipeline cache mutexes (#19195 )

2026-02-01 18:47:29 -08:00

ggml-zdnn

ggml-zdnn : mark zDNN buffers as non-host (#18967 )

2026-01-22 01:16:21 +01:00

ggml-zendnn

ggml-zendnn : resolve ZenDNN backend cross-module symbol dependency (#19159 )

2026-01-29 12:28:57 +08:00

CMakeLists.txt

hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150 )

2026-01-29 12:33:21 -08:00

ggml-alloc.c

llama: automatically set parameters not set by the user in such a way that maximizes GPU utilization (#16653 )

2025-12-15 09:24:59 +01:00

ggml-backend-dl.cpp

hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150 )

2026-01-29 12:33:21 -08:00

ggml-backend-dl.h

hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150 )

2026-01-29 12:33:21 -08:00

ggml-backend-impl.h

llama: use host memory if device reports 0 memory (#18587 )

2026-01-09 05:34:56 +08:00

ggml-backend-reg.cpp

hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150 )

2026-01-29 12:33:21 -08:00

ggml-backend.cpp

ggml-backend: fix async set/get fallback sync (#19179 )

2026-02-02 10:00:05 +01:00

ggml-common.h

llama : add gpt-oss (#15091 )

2025-08-05 22:10:36 +03:00

ggml-impl.h

ggml : add ggml_build_forward_select (#18550 )

2026-01-19 20:03:19 +02:00

ggml-opt.cpp

finetune: SGD optimizer, more CLI args (#13873 )

2025-08-14 12:03:57 +02:00

ggml-quants.c

ggml : fix uninitialized is_on_grid in quantize_row_iq3_xxs_impl (#15928 )

2025-09-23 10:25:20 +02:00

ggml-quants.h

llama : add gpt-oss (#15091 )

2025-08-05 22:10:36 +03:00

ggml-threading.cpp

ggml : build backends as libraries (#10256 )

2024-11-14 18:04:35 +01:00

ggml-threading.h

remove CMAKE_WINDOWS_EXPORT_ALL_SYMBOLS (#10797 )

2024-12-12 19:02:49 +01:00

ggml.c

ggml: added cleanups in ggml_quantize_free (#19278 )

2026-02-03 08:43:39 +02:00

ggml.cpp

ggml : Print backtrace on uncaught C++ exceptions (ggml/1232)

2025-06-01 13:43:57 +03:00

gguf.cpp

GGUF: check that tensor size is representable (#19072 )

2026-01-24 21:57:51 +01:00