TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

History

qixiang-99 b165f8bc97 fix/improve kvcache allocation in PyTorch runtime (#5933 ) Signed-off-by: qixiang-99 <203170375+qixiang-99@users.noreply.github.com>		2025-08-26 12:40:22 +08:00
..
bindings.cpp	[TRTLLM-6881][feat] Include attention dp rank info with KV cache events (#6563 )	2025-08-07 14:17:07 +02:00
bindings.h	fix: Move all casters to customCasters. (#3945 )	2025-05-02 19:08:28 +08:00
executor.cpp	Update TensorRT-LLM (#2436 )	2024-11-12 15:27:49 +08:00
executor.h	Update TensorRT-LLM (#2562 )	2024-12-11 00:31:05 -08:00
executorConfig.cpp	fix/improve kvcache allocation in PyTorch runtime (#5933 )	2025-08-26 12:40:22 +08:00
executorConfig.h	fix: Move all casters to customCasters. (#3945 )	2025-05-02 19:08:28 +08:00
request.cpp	[None][fix] acceptance rate calculation fix in benchmark_serving (#6746 )	2025-08-19 17:29:36 +08:00
request.h	fix: Move all casters to customCasters. (#3945 )	2025-05-02 19:08:28 +08:00