TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

History

qixiang-99 b165f8bc97 fix/improve kvcache allocation in PyTorch runtime (#5933 ) Signed-off-by: qixiang-99 <203170375+qixiang-99@users.noreply.github.com>		2025-08-26 12:40:22 +08:00
..
cacheCommunicator.h	Agent interface impl for NIXL (#4125 )	2025-05-22 09:09:41 +08:00
dataTransceiverState.h	[None][chore] No-op changes to support context parallelism in disaggregated serving later (#7063 )	2025-08-21 08:21:27 -07:00
disaggServerUtil.h	Update TensorRT-LLM (#2792 )	2025-02-18 21:27:39 +08:00
executor.h	fix/improve kvcache allocation in PyTorch runtime (#5933 )	2025-08-26 12:40:22 +08:00
serialization.h	[TRTLLM-6881][feat] Include attention dp rank info with KV cache events (#6563 )	2025-08-07 14:17:07 +02:00
tensor.h	Update TensorRT-LLM (#1918 )	2024-07-09 14:42:22 +08:00
transferAgent.h	Agent interface impl for NIXL (#4125 )	2025-05-22 09:09:41 +08:00
types.h	[TRTLLM-5000][feat] NGrams V2 (#4569 )	2025-06-27 23:00:17 +08:00