TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

History

HuiGao-NV 3ade9375ba feat: Run PyExecutor's inference flow to estimate max_num_tokens for kv_cache_manager (#3092 ) Signed-off-by: Hui Gao <huig@nvidia.com>		2025-04-10 18:29:40 +08:00
..
tensorrt_llm	feat: Run PyExecutor's inference flow to estimate max_num_tokens for kv_cache_manager (#3092 )	2025-04-10 18:29:40 +08:00