TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

History

HuiGao-NV 3ade9375ba feat: Run PyExecutor's inference flow to estimate max_num_tokens for kv_cache_manager (#3092 ) Signed-off-by: Hui Gao <huig@nvidia.com>		2025-04-10 18:29:40 +08:00
..
batch_manager	feat: Run PyExecutor's inference flow to estimate max_num_tokens for kv_cache_manager (#3092 )	2025-04-10 18:29:40 +08:00
common	Feat: Variable-Beam-Width-Search (VBWS) part3 (#3338 )	2025-04-08 23:51:27 +08:00
deep_gemm	fix: remove DeepGEMM line info (#3411 )	2025-04-09 18:01:02 +08:00
executor	Feat: Variable-Beam-Width-Search (VBWS) part3 (#3338 )	2025-04-08 23:51:27 +08:00
kernels	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
layers	v1.2 (#3082 )	2025-03-26 23:31:29 +08:00
plugins/api	Update TensorRT-LLM (#2532 )	2024-12-04 21:16:56 +08:00
runtime	Feat: Variable-Beam-Width-Search (VBWS) part3 (#3338 )	2025-04-08 23:51:27 +08:00