TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

History

Arthur Rasmusson 812b1abf86 feature: KV Cache GPUDirect Storage (#3209 ) Signed-off-by: Arthur Rasmusson <47877520+arthurrasmusson@users.noreply.github.com.> Co-authored-by: Robin Kobus <19427718+Funatiq@users.noreply.github.com> Co-authored-by: Aurelien Chartier <2567591+achartier@users.noreply.github.com>		2025-05-28 23:27:43 +00:00
..
bindings.cpp	[TRTLLM-5000][feat] Pytorch implementation of ngram drafter (#3936 )	2025-05-21 10:40:00 +08:00
bindings.h	fix: Move all casters to customCasters. (#3945 )	2025-05-02 19:08:28 +08:00
executor.cpp	Update TensorRT-LLM (#2436 )	2024-11-12 15:27:49 +08:00
executor.h	Update TensorRT-LLM (#2562 )	2024-12-11 00:31:05 -08:00
executorConfig.cpp	Feat: Variable-Beam-Width-Search (VBWS) part4 (#3979 )	2025-05-12 22:32:29 +02:00
executorConfig.h	fix: Move all casters to customCasters. (#3945 )	2025-05-02 19:08:28 +08:00
request.cpp	feature: KV Cache GPUDirect Storage (#3209 )	2025-05-28 23:27:43 +00:00
request.h	fix: Move all casters to customCasters. (#3945 )	2025-05-02 19:08:28 +08:00