llama.cpp

mirror of https://github.com/ggml-org/llama.cpp.git synced 2026-07-02 01:00:20 +00:00

Files

T

Ruben Ortlam 0f7fada56b cuda: reset cuda context after reading memory size (#23935 )

* cuda: reset device in get_memory function if no backend is active

* also count device and host buffers

* exclude hip and musa from counting and device reset

* use device mutex instead of atomic

* undo backend_free function move

2026-06-08 10:22:44 +02:00

cmake

ggml : Parallelize quant LUT init (#23595 )

2026-05-25 10:15:46 +03:00

include

TP: quantized KV cache support (#23792 )

2026-06-01 12:30:10 +02:00

src

cuda: reset cuda context after reading memory size (#23935 )

2026-06-08 10:22:44 +02:00

.gitignore

vulkan : cmake integration (#8119 )

2024-07-13 18:12:39 +02:00

CMakeLists.txt

ggml : bump version to 0.13.1 (ggml/1523)

2026-05-29 09:56:08 +03:00