TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

History

Yueh-Ting (eop) Chen 199f306984 [None][chore][kv cache manager] Dead code elimination, we no longer record/fetch through WindowBlockManager:: mContextBlocksByHash (#6249 ) No functional change is intended in this MR. `WindowBlockManager::mCachedBlocksRoot` is now who is responsible for the bookkeeping of the `KVCacheBlock`, and the `mNextBlocks` is now the actual hash map that fetches the block. The `mEnableHashKey` knob and related hashing is removed. Signed-off-by: eopXD <yuehtingc@nvidia.com>		2025-08-10 09:10:10 -04:00
..
batch_manager	[None][chore][kv cache manager] Dead code elimination, we no longer record/fetch through WindowBlockManager:: mContextBlocksByHash (#6249 )	2025-08-10 09:10:10 -04:00
common	[None] [feat] Add model gpt-oss (#6645 )	2025-08-07 03:04:18 -04:00
deep_gemm	fix: fix license bug (#5200 )	2025-06-13 18:58:15 +08:00
executor	[TRTLLM-6881][feat] Include attention dp rank info with KV cache events (#6563 )	2025-08-07 14:17:07 +02:00
kernels	fix: compatibility with CUDA < 12.9 on `__CUDA_ARCH_SPECIFIC__` macro (#5917 )	2025-07-28 16:02:26 +08:00
layers	v1.2 (#3082 )	2025-03-26 23:31:29 +08:00
plugins/api	Update TensorRT-LLM (#2532 )	2024-12-04 21:16:56 +08:00
runtime	[TRTLLM-6785][feat] BREAKING CHANGE Enable TRTLLM sampler by default (#6216 )	2025-08-07 22:19:37 -04:00