llama.cpp

mirror of https://github.com/ggml-org/llama.cpp.git synced 2026-06-28 07:10:21 +00:00

Author	SHA1	Message	Date
Aman Gupta	f8f0a47a55	cuda: reserve space for quantize kv-cache at startup (#23907 ) * cuda: reserve space for quantize kv-cache at startup * address review comments * remove forward decl Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * remove assert in ggml-cuda.cu Co-authored-by: Johannes Gäßler <johannesg@5d6.de> --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de> b9489	2026-06-03 18:39:59 +08:00
Georgi Gerganov	06938ac129	tests : add support for qwen3 SSM archs (#24031 ) * tests : add support for qwen3 SSM archs * arch : add LLM_KV_ATTENTION_RECURRENT_LAYERS * cont : naming + TODOs b9488	2026-06-03 10:15:27 +03:00
Alessandro de Oliveira Faria (A.K.A.CABELO)	d545a2a993	update BoringSSL to 0.20260526.0 (#23794 ) b9487	2026-06-03 07:42:58 +02:00
Georgi Gerganov	4da6370d43	ci : disable ccache for msvc windows release jobs (#23911 ) b9486	2026-06-03 08:05:21 +03:00
Ryan Mangeno	e3666269f9	arg : removed unecesary mmproj download when users pass --no-mmproj (#23425 ) b9485	2026-06-03 08:04:46 +03:00
lhez	63e66fdd23	opencl: use flat variants of q4_K and q6_K gemv for very large M (#24006 ) b9484	2026-06-02 14:16:17 -07:00
Max Krasnyansky	5c394fdc8b	hexagon: profiler output fix and script updates (#24042 ) * hex-ops: fix profiler output (ie remove the redundant NONEs) * hex-prof: update profiling script to support tot.usec column b9483	2026-06-02 14:08:29 -07:00
Mikhail Podvitskii	4fb16eccce	model: add Mellum architecture (#23966 ) * model: support for Mellum architecture * model: improve mellum.py formatting * model: improve mellum.py formatting once again * deps: downgrade transformers to 4.57.6 (to fix CI) * deps: remove huggingface_hub dependency * deps: remove huggingface_hub from test requirements --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> b9482	2026-06-02 22:11:12 +03:00
Hans Florian	bfb4308b05	model : support granite multilingual embeddings R2 (ibm-granite/granite-embedding-{97,311}m-multilingual-r2) (#22716 ) * Add support for the ibm-granite/granite-embedding-{97m,311m}-multilingual-r2 embedding models: * Added a version of the gpt4o tokenizer that has a fixed regex (better handling of marks), and different token merging setting for the 97m model * Reused gemma4 tokenizer for the 311m model * granite-embedding--multilingual-r2 : add support SwiGLU FFN for Granite Embedding Multilingual R2 added new GGUF key <arch>.hidden_activation (LLM_KV_HIDDEN_ACT) + writer * added a forward declaration of llm_ffn_op_type to llama-hparams.h * added llm_ffn_op in hparams * added LLM_FFN_NONE = 0 sentinel to llm_ffn_op_type (value-initialization), modern-bert: explicitly assigns LLM_FFN_GEGLU before reading GGUF (unchanged). * centralized hidden_act mapping in llama-model.cpp, added llm_ffn_op_type_from_string() helper, mirroring rope_scaling_type/llama_rope_scaling_type_from_string() * modern-bert reads the GGUF key (when present) and uses the resulting op in its FFN graph * Added granite-embedding-{97m,311m}-multilingual-r2 to the converter code * Added the hashes for the granite embedding multilingual R2 models * Set the hidden_activation in the GGUF if the field is present in config.json (such as for the granite embedding models) b9481	2026-06-02 17:55:11 +02:00
Piotr Wilkin (ilintar)	2187e00337	StepFun 3.5 MTP (#23274 ) * StepFun 3.5 MTP * Simplify to single layer * Rollback core changes * fix flake8 errors * Remove scripts * modify to convention * Apply suggestions from code review Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * dos2unix --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> b9480	2026-06-02 17:44:35 +02:00
Daniel Bevenius	0b7154066e	common : fix state save in common_prompt_batch_decode (#23468 ) * common : fix state save in common_prompt_batch_decode This commit addresses a bug in common_prompt_batch_decode that affects the session state store/restore in completion.cpp and save-load-state.cpp. The motivation for this is that currently the code is saving n-1 tokens in both the session_tokens and in the KV cache. Then when loading the session tokens, and if the prompt matches, it would replay the last saved token (n-1) into the next position, effectively replaying the same token in the wrong position. The fix is to store all n tokens in session_tokens, while the memory state only reflects n-1 processed tokens as the saving happens before the last token is decoded in common_prompt_batch_decode. I ran both completion.cpp and save-load-state.cpp with a transformer, a recurrent, and a hybrid model. Resolves: https://github.com/ggml-org/llama.cpp/issues/23400 Co-authored-by: fairydreaming <166155368+fairydreaming@users.noreply.github.com> b9479	2026-06-02 15:44:15 +02:00
Xuan-Son Nguyen	60130d18f9	server: add SSE ping interval (#24013 ) b9478	2026-06-02 14:14:55 +02:00
Georgi Gerganov	a468b89018	ci : reduce self-hosted server workflow jobs (#24012 ) Reduce the number of parallel jobs in server-self-hosted.yml by stacking test configurations as sequential steps within a single job, following the pattern from #23927. - server-metal: 4 matrix jobs -> 1 job with 4 sequential test steps - server-cuda: 2 matrix jobs -> 1 job with 2 sequential test steps - server-kleidiai: removed unnecessary single-entry matrix - removed unused Setup Node.js step from server-metal Total: 7 parallel jobs -> 3 parallel jobs Assisted-by: llama.cpp:local pi	2026-06-02 13:17:59 +03:00
Mikhail Podvitskii	d5ab0834ab	docs : update HOWTO-add-model.md (#23883 ) * docs: update HOWTO-add-model.md with new model registration and graph-building instructions * docs: improve formatting in HOWTO-add-model.md * Update docs/development/HOWTO-add-model.md Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>	2026-06-02 11:40:22 +02:00
Marcos Del Sol Vives	69cea5b669	ui: simplify network error handling (#23431 ) Previously error to string conversion was split in two different files, with one converting errors into strings, and another function analyzing those strings to generate yet another string. Now the the error handling for network fetches has been centralised and uses directly HTTP error codes whereas possible to generate the human-readable error strings. It also fixes an issue where all JSON errors reported from the backend, such as "Invalid API key", would get turned incorrectly in to "Failed to connect to server" due to poor matching logic in the now-gone getErrorMessage function.	2026-06-02 10:45:25 +02:00
Aleksander Grygier	f8e67fc583	ui: Add Thinking mode toggle with reasoning effort levels + improvements for Chat Form Add Action UI (#23434 ) * feat: Add "Thinking" toggle and status icon + redesign Chat Form Actions Add panel * test: Update test reference * fix: Icon * fix: E2E test command * fix: wait for greeting h1 to be visible in e2e test * fix: remove duplicate PDF option in attachment dropdown * fix: use label-based group toggle to avoid stale references * refactor: inline MCP server and tool toggles in mobile sheet * fix: serve correct build directory in e2e playwright config * feat: add reasoning effort levels selector in model dropdown * feat: Reasoning effort * refactor: Make server origin configurable via environment variable * feat: Add chat template thinking detector utility * feat: Add thinking support detection to models store * refactor: Update model selector components with thinking detection and message-specific indicators * feat: Update chat form components for model selection and thinking support * feat: Improve Reasoning controls UI * refactor: Apply suggestions from code review Co-authored-by: Aleksander Grygier <aleksander.grygier@gmail.com> * fix: Model tags * refactor: Cleanup * refactor: Remove unneeded components * refactor: Cleanup b9474	2026-06-02 10:23:19 +02:00
Georgi Gerganov	2365315955	kv-cache : SWA checkpoints store only non-masked cells (#23981 ) b9473	2026-06-02 11:06:29 +03:00
forforever73	f7a0777a5c	convert : support Step3.7-Flash (#23845 ) * feat: support step3.7 * fix: register Step-3.7 BPE pre-tokenizer hash * delete fromjson * register step3.7 arch to Step35Model * drop vit projector in base filter * Apply suggestion from @CISC Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * restore blank line --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>	2026-06-02 09:54:49 +02:00
Georgi Gerganov	4f3a4beb8d	llama : deprecate `llama_set_warmup` (#24009 ) * llama : deprecate `llama_set_warmup` * cont : fix type Co-authored-by: Daniel Bevenius <daniel.bevenius@gmail.com> --------- Co-authored-by: Daniel Bevenius <daniel.bevenius@gmail.com> b9471	2026-06-02 10:30:38 +03:00
Max Krasnyansky	8f7f3bf141	hexagon: MUL_MAT, MUL_MAT_ID, FLASH_ATTN and GDN cleanup and optimizations for latest models (#23989 ) * hex-mm: initial support for F32 * F32 -> F32 matmuls * hex-rms-norm: fix src1 stride use in fused rms_norm_mul * hex-ops: clear spad pointers in the ops that clober it This fixes an odd case where fused rms-norm-mul was failing but only in qwen3.5-2B and only at searth op-bath sizes. * hmx-mm: add support for F32 * F32 -> F32 matmul_2d on HMX Decided to use Q4_0 * F32 -> F32 matmul for this. Q4_0 gets dequantized and tiled into F16, and here we quantize and tile F32 into F16. Super simple and pretty efficient. * hmx-mm: route f16 2D matmuls through the same kernel used for all other types * hmx-mm: re-introduce pipelined vs non-pipelined mode that we used to have but is much more generic way This update futher improves matmul performance and at the same time removes most of the redudant logic we had in different paths. * hmx-fa: slighlty improved pipeline simimar to matmul updates * hmx-mm: initial version of MAT_MUL_ID support for HMX * hmx-mm: fixed mxfp4 handling for MUL_MAT_ID * hex-gdn: optimize GATED_DELTA_NET DMA prefetch/double-buff, vectorize everything with HVX, in other words -- the usual :) * hmx-mm: missed one more case where we can use fastmod * hexagon: update DCVS settings for a slight perf bump * hmx-fa: use fastdiv in hmx-flash-attn * hmx-fa: precompute slope values to avoid disrupting the inner loop * hvx-utils/fa: new HVX helpers for powf and logf and using those to speed up FA alibi * hex-ops: fixed a bug in fusion logic that was messing up the order of the src tensors when some srcs are empty * hex-fa: correctly fallback to HVX if we have sinks or the dims are not quite right b9470	2026-06-01 23:40:08 -07:00
Todor Boinovski	d178a11818	hexagon: add gelu_quick (#24007 ) b9469	2026-06-01 23:19:07 -07:00
Pascal	354ebac8cb	server: real-time reasoning interruption via control endpoint (#23971 ) * server: real-time reasoning interruption via control endpoint Builds on the manual reasoning budget trigger from #23949. Adds a CONTROL task that mirrors the CANCEL path on the live slot and calls common_sampler_reasoning_budget_force to end thinking mid-generation. POST /v1/chat/completions/control with { id_slot, action }, opt-in reasoning_control arms the budget sampler on demand. Router and single model. Minimal WebUI button as a skeleton for further UI work. * ui: track reasoning phase via explicit streaming state Add isReasoning to the chat store, mirroring the isLoading pattern: per conversation map, private setter, public accessor and reactive export. Set from the stream callbacks, true on reasoning chunks, false on the first content chunk, reset on stream end and resynced on conversation switch. The skip button now keys off isReasoning so it shows only during the thinking phase, not the whole generation. * ui: extract control endpoint and action into constants Move the chat completion routes, the slots route and the reasoning control action out of chat.service into api-endpoints and a dedicated control-actions module. No behavior change, drops the magic strings so the control protocol has a single source of truth. * server: target reasoning control by completion id Address @ngxson review on the control endpoint. Switch from id_slot to the chat completion id to avoid a TOCTOU: the slot can be reassigned between the lookup and the control request, so matching the live completion (oaicompat_cmpl_id) is safe and a finished one simply matches nothing. Rename the action to reasoning_end, guard it on the reasoning_control flag of the target slot, and reduce the response to {success} with an optional message. * ui: target reasoning control by completion id Keep the streamed completion id on the message and post it back to the control endpoint instead of probing /slots. Drops the slot discovery and the TOCTOU that came with it. Action renamed to reasoning_end, response read as {success}. * server: address review from @ngxson Move the control fields into task_params and drop the redundant comments on the control path. * server: document the reasoning control endpoint * Update tools/ui/src/lib/types/database.d.ts Co-authored-by: Aleksander Grygier <aleksander.grygier@gmail.com> * ui: rename cmplId to completionId Per @allozaur review, clearer name for the streamed completion id. * ui: wire completion id capture through the agentic flow The webui streams through the agentic flow, which relayed onModel but not onCompletionId, so the completion id never reached the message and the control request was never sent. Relay it through the flow and its callbacks type, declare id on the chunk type, and log an explicit error when the button fires without a usable id. * ui: target reasoning control model from the message The model is a property of the completion, so read it from the streaming message like the id, not from the model dropdown which is unrelated UI state. Makes the request self-consistent by construction instead of just unlikely to drift. --------- Co-authored-by: Aleksander Grygier <aleksander.grygier@gmail.com> b9468	2026-06-02 07:26:20 +02:00
Anav Prasad	1fd5f48037	clean up unused variables warnings (#23975 ) b9467	2026-06-02 10:38:37 +08:00
lhez	210a6570ce	opencl: fix compiler warnings for non-adreno path (#23922 ) * opencl: fix compiler warnings for non-adreno path * opencl: fix const cast warning b9466	2026-06-01 19:15:09 -07:00
Masashi Yoshimura	b8275a8acc	revert to using global_invocation_id for cpy shader (#23955 )	2026-06-01 16:59:06 -07:00
Georgi Gerganov	5dcb711666	speculative : fix n_outputs_max and remove draft-simple auto-enable (#23988 ) * speculative : add common_speculative_n_max helper function Extract the speculative max-draft-size logic from server_n_outputs_max into a reusable common_speculative_n_max() function in common/speculative. Assisted-by: llama.cpp:local pi * cont : draft context always has n_parallel outputs * llama : log n_outputs_max * speculative : remove draft-simple auto-enable * ci : enable server tests on PRs b9464	2026-06-01 22:26:58 +03:00
Christian Hoener zu Siederdissen	5aa3a64596	nix : add nix-nodejs facilities to build Web UI (#23846 ) * nix: add nix-nodejs facilities to build Web UI Build the Web UI locally using standard Nix systems for building NodeJS packages. - Create derivation for the web UI - npm dependencies are imported via buildNodeModules. Does not require setting any shasum. - Copy build artifacts to the correct folders. - Prevents having to download from huggingface.co Fixes #23067 * nix: simplify webui derivation using LLAMA_UI_OUT_DIR - Move npm build to installPhase with LLAMA_UI_OUT_DIR=$out to write output directly to the Nix store - Copy built assets to tools/ui/dist (source tree) instead of build/tools/ui/dist so CMake's copy_src_dist() finds them	2026-06-01 14:01:26 -04:00
shaofeiqi	27d9ed8397	opencl: add basic support for q5_0 and q5_1 (#23548 ) * opencl: add general q5_0 support * opencl: add general q5_1 support * opencl: support non-uniform workgrp size --------- Co-authored-by: Li He <lih@qti.qualcomm.com>	2026-06-01 10:06:50 -07:00
Adrien Gallouët	335abed17d	vendor : update cpp-httplib to 0.46.1 (#23980 ) Signed-off-by: Adrien Gallouët <angt@huggingface.co>	2026-06-01 19:40:10 +03:00
Aman Gupta	de6f727aae	llama: limit max outputs of `llama_context` (#23861 ) * llama: save more VRAM by reserving n_outputs == n_seqs when possible * add n_outputs_per_seq * move n_outputs_max to server-context * change ubatch to batch everywhere b9460	2026-06-01 18:01:38 +03:00
Shrivas Shankar	95b8b8ec1a	metal: template GLU kernels to support f16/f32 (#23882 ) Drops the hardcoded f32 GLU kernels in favor of a single template. We now load/store in the native tensor type (half or float) to save memory bandwidth, but keep the actual ALU compute in float to avoid exploding math in geglu/swiglu. Also opened up the dispatch gate to allow f16 inputs. b9459	2026-06-01 15:40:28 +03:00
Jeff Bolz	55ac0909e5	vulkan: don't hold the device mutex while compiling pipelines (#23641 ) * vulkan: don't hold the device mutex while compiling pipelines We need to hold a lock while we traverse all pipelines and lazily initialize them, but we don't need to hold it while the pipeline is being compiled. And it doesn't need to be the same lock as the device mutex. We call load_shaders each time a pipeline is needed, so we only need to compile that one pipeline (and, for example, don't want to end up compiling a pipeline that another thread should be compiling). * remove 'needed' b9458	2026-06-01 14:04:01 +02:00
Winston Ma	bef69f1306	vulkan: reduce host memory lock contention (#23376 ) * vulkan: reduces lock contention * replace unique_lock with lock_guard b9457	2026-06-01 14:03:32 +02:00
o7si	5aba5364d9	vocab: add normalizer.lowercase support to WPM (#23899 ) * vocab : add jina-embeddings-v2-base-zh (whitespace tokenizer) * vocab : add normalizer.lowercase support to WPM * vocab : default normalizer.lowercase to false for whitespace pre-tokenizer	2026-06-01 14:26:47 +03:00
Johannes Gäßler	8e6fff84de	TP: quantized KV cache support (#23792 ) * TP: quantized KV cache support * fix partial view * remove overly strict assert b9455	2026-06-01 12:30:10 +02:00
Georgi Gerganov	02a57017f6	security : disable private disclosures (#23963 )	2026-06-01 13:14:12 +03:00
Junwon Hwang	48b88c3b00	model: Add EXAONE 4.5 implementations (#21733 ) * Add EXAONE 4.5 and Add GQA for MMproj * mtmd: EXAONE 4.5 vision markers and projector path EXAONE 4.5 uses <vision> and </vision> for image boundaries; Qwen keeps <\|vision_start\|> and <\|vision_end\|>. Route EXAONE 4.5 through the Qwen2.5-VL-style encode path (window attention pattern, optional mmproj input norm). Update exaone4_5 projector weights and convert_hf_to_gguf for mmproj export. * mtmd: load EXAONE4 nextn tensors correctly Align EXAONE4 tensor registration with EXAONE_MOE for NextN/MTP slots and avoid skip-flag propagation on duplicated rope_freqs so model loading succeeds for EXAONE 4.5 GGUF. * Minor fixes * Address PR feedback * Address PR feedback * Fix EXAONE after merge * Fix EXAONE 4.5 conversion * Address PR feedback * Refactor EXAONE 4.5 conversion * Address PR feedback * Fix unintended deletion * Minor fix --------- Co-authored-by: LG-AI-EXAONE <exaonemodels@lgresearch.ai> b9453	2026-06-01 11:48:53 +02:00
Matt Corallo	19620004f5	vulkan: Block-load Q3_K/Q6_K block data and subtract on 32b ints (#23056 ) Q2_K/Q3_K/Q6_K do much better when using MMVQ on Intel BMG even though they're only 2-byte aligned, and Q3_K still wins on NVIDIA as well. mesa isn't all that great at coalescing back-to-back loads from alternating arrays, so we force it instead. Further, we can do subtraction directly on a full int32_t rather than an i8vec4 with bit twiddling because the high bit is always free to start. On Intel BMG on mesa, the switch to MMVQ provides an immediate ~57% perf increase in tg128 for unsloth/Qwen3.5-9B-GGUF:Q3_K and ~78% perf increase in tg128 for unsloth/Qwen3.5-9B-GGUF:Q6_K. The futher switch to block loads leads to a ~24% perf increase in tg128 for unsloth/Qwen3.5-9B-GGUF:Q3_K and a ~48% perf increase in tg128 for unsloth/Qwen3.5-9B-GGUF:Q6_K. Finally, Xe2 wins on MMVQ even for small k, so we take the NVIDIA override for K quants on Xe2 as well. b9452	2026-06-01 11:46:48 +02:00
Winston Ma	f8c0a19d46	vulkan: Removed unused functions (#23175 ) b9451	2026-06-01 11:46:23 +02:00
Aldehir Rojas	5254a7994d	common : support manually triggering the reasoning budget end sequence (#23949 )	2026-06-01 11:37:11 +02:00
Georgi Gerganov	e22b0de60d	ci : add missing Linux label to cpu-x64-high-perf runner (#23958 ) Fixes: https://github.com/ggml-org/llama.cpp/pull/23927#discussion_r3332213086 The cpu-x64-high-perf job was missing the Linux label in its runs-on specification, causing the runner to not be discovered. All other self-hosted Linux jobs include this label. Assisted-by: llama.cpp:local pi	2026-06-01 10:39:59 +03:00
Neo Zhang	a51142497a	[SYCL] Support Q4_1, Q5_0, Q5_1 in Flash-attention (#23812 ) * support Q4_1, Q5_0, Q5_1 * update ut case	2026-06-01 09:53:53 +03:00
Neo Zhang	4162522688	[SYCL] Add more types in GET_ROWS OP (#23710 ) * add to support Q1_0, NVFP4, IQ2_XXS, IQ2_XS, IQ2_S, IQ3_XXS, IQ1_S, IQ1_M, IQ3_S, IQ4_NL, IQ4_XS, I32, MXFP4, Q2_K, Q3_K, Q5_K, and Q6_K in GET_ROWS OP * correct the link	2026-06-01 09:53:04 +03:00
Neo Zhang	44e211cecf	sycl : Optimize Q3_K mul_mat by reorder (#23725 )	2026-06-01 09:50:55 +03:00
Eve	af6528e6df	ci: remove redundant or duplicate jobs (#23927 ) * remove redundant apple job openvino gpu and cpu test can share the same build and machine Update build-rpc.yml Update build-openvino.yml cpu any doesnt make sense as we have an arm job already, so do high perf on both x86 and arm remove duplicate x86 vulkan combine backend sampling Update server.yml run server on arm as windows is x86 * emdawn on one machine only * fix openvino, remove cpu tag as we dont have many x64 machines with that tag b9445	2026-06-01 06:32:17 +03:00
Eric Zhang	6f165c1c64	server : handle If-None-Match weak ETags (#23916 ) b9444	2026-05-31 16:21:08 -05:00
Georgi Gerganov	399739d5c5	ci : limit trigger paths for the CPU workflow (#23938 )	2026-05-31 19:02:47 +03:00
o7si	d4c8e2c29c	vocab : add tokenizer support for jina-embeddings-v2-base-zh (#18756 ) * vocab : add jina-embeddings-v2-base-zh (whitespace tokenizer) * lowercase defaults to true * type fix --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> b9442	2026-05-31 12:37:35 +02:00
Eric Zhang	3292da09f6	ui: fix ETag truncation with MSVC compiler (#23917 ) b9441	2026-05-31 11:21:23 +02:00
Vladislav	e6123e2080	docs : update ZenDNN docs for Q8 support (#23791 ) * docs zendnn added information about Q8 support * docs zendnn rm unnecessary data * docs update, links to ZenDNN docs provided * docs zenDNN update: clarified explanation * docs zenDNN update: one more explanation clarified --------- Co-authored-by: plotnikov.v10 <plotnikov.v10@wb.ru>	2026-05-31 10:26:42 +02:00

1 2 3 4 5 ...

9489 Commits