TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

History

Pengyun Lin 039f7e3118 [https://nvbugspro.nvidia.com/bug/5243740 ][fix] deduce default max_tokens for trtllm-serve (#4265 ) * Deduce default max_tokens for trtllm-serve Signed-off-by: Pengyun Lin <81065165+LinPoly@users.noreply.github.com> * Improve executor_config.max_seq_len assignment in TRT workflow Signed-off-by: Pengyun Lin <81065165+LinPoly@users.noreply.github.com> * Enhance error message Signed-off-by: Pengyun Lin <81065165+LinPoly@users.noreply.github.com> * Add deduced max_tokens test Signed-off-by: Pengyun Lin <81065165+LinPoly@users.noreply.github.com> --------- Signed-off-by: Pengyun Lin <81065165+LinPoly@users.noreply.github.com>		2025-05-19 00:34:40 +08:00
..
_torch	perf: Eliminate the need for attention DP padding when possible (#3439 )	2025-05-17 13:30:55 +08:00
api_stability	add changes for fp8, nemotron-nas, API (#4180 )	2025-05-18 23:27:25 +08:00
bindings	refactor: use x is None instead of x == None. (#4244 )	2025-05-15 20:00:04 +08:00
disaggregated	feat: add kv cache aware router (#3831 )	2025-05-12 07:23:57 -04:00
llmapi	[https://nvbugspro.nvidia.com/bug/5243740 ][fix] deduce default max_tokens for trtllm-serve (#4265 )	2025-05-19 00:34:40 +08:00
others	test: reorganize tests folder hierarchy (#2996 )	2025-03-27 12:07:53 +08:00
scaffolding	feat: support benchmark on scaffolding (#3328 ) (#4286 )	2025-05-16 12:28:49 +08:00
tools	Update TensorRT-LLM (#2936 )	2025-03-18 21:25:19 +08:00
trt	[TRTLLM-3330][feat] Support DeepSeek-R1 W4A8 on Hopper (#4123 )	2025-05-14 15:48:07 +08:00
utils	add changes for fp8, nemotron-nas, API (#4180 )	2025-05-18 23:27:25 +08:00
conftest.py	Add thread leak check and fix thread/memory leak issues. (#3270 )	2025-04-08 19:03:18 +08:00
dump_checkpoint_stats.py	Update TensorRT-LLM (#2936 )	2025-03-18 21:25:19 +08:00
profile_utils.py	Update TensorRT-LLM (#2936 )	2025-03-18 21:25:19 +08:00
pytest.ini	move the reset models into `examples/models/core` directory (#3555 )	2025-04-19 20:48:59 -07:00
test_model_runner_cpp.py	Update TensorRT-LLM (#2936 )	2025-03-18 21:25:19 +08:00
test_pip_install.py	relax the limitation of setuptools (#2992 )	2025-03-24 13:36:10 +08:00