TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

History

Michal Guzek 08d57123f9 [nvbug/5374773] chore: Add a runtime flag to enable fail fast when attn window is too large to fit at least one sequence in KV cache (#5974 ) Signed-off-by: moraxu <mguzek@nvidia.com>		2025-07-25 18:10:40 -04:00
..
bindings.cpp	feat: nanobind bindings (#6185 )	2025-07-21 08:56:57 +01:00
bindings.h	fix: Move all casters to customCasters. (#3945 )	2025-05-02 19:08:28 +08:00
executor.cpp	Update TensorRT-LLM (#2436 )	2024-11-12 15:27:49 +08:00
executor.h	Update TensorRT-LLM (#2562 )	2024-12-11 00:31:05 -08:00
executorConfig.cpp	[nvbug/5374773] chore: Add a runtime flag to enable fail fast when attn window is too large to fit at least one sequence in KV cache (#5974 )	2025-07-25 18:10:40 -04:00
executorConfig.h	fix: Move all casters to customCasters. (#3945 )	2025-05-02 19:08:28 +08:00
request.cpp	[TRTLLM-6104] feat: add request_perf_metrics to LLMAPI (#5497 )	2025-06-27 17:03:05 +02:00
request.h	fix: Move all casters to customCasters. (#3945 )	2025-05-02 19:08:28 +08:00