TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

Author	SHA1	Message	Date
QI JUN	616d1df7a0	[None][chore] set the default value of max_num_tokens explicitly (#8208 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>	2025-10-14 23:03:02 -07:00
sychen52	6a6124dcb5	[OMNIML-2336][feat] w4a8 nvfp4 fp8 exports scale factor properly (#8180 ) Signed-off-by: Shiyang Chen <shiychen@nvidia.com> Co-authored-by: Shiyang Chen <shiychen@omniml-a6.nvidia.com>	2025-10-15 13:41:27 +08:00
Jin Li	206a9930df	[https://nvbugs/5547435 ][fix] Fix a merge conflict (#8365 ) Signed-off-by: Jin Li <59594262+liji-nv@users.noreply.github.com>	2025-10-15 10:43:10 +08:00
Emma Qiao	493da020c1	[TRTLLM-7351][infra] Add isolate marker for L0 (#7497 ) Signed-off-by: qqiao <qqiao@nvidia.com> Signed-off-by: Emma Qiao <qqiao@nvidia.com> Co-authored-by: Yanchao Lu <yanchaol@nvidia.com>	2025-10-14 16:58:14 -07:00
dongfengy	9d855f47ad	[None][fix] Remove outdated test waives for GPTOSS (#8183 ) Signed-off-by: Dongfeng Yu <dongfengy@nvidia.com>	2025-10-14 16:20:38 -07:00
Lizhi Zhou	22471ecc67	[TRTLLM-7846][feat] implement etcd storage for disagg cluster (#8210 ) Signed-off-by: Lizhi Zhou <1432185+reasonsolo@users.noreply.github.com>	2025-10-14 16:48:41 -04:00
Michal Guzek	1cdb0b62c3	[https://nvbugs/5563469 ][fix] Temporarily disable test_nemotron_nano_8b_lora_torch in L0 due to Torch non-determinism (#8206 ) Signed-off-by: Michal Guzek <mguzek@nvidia.com>	2025-10-14 17:55:28 +02:00
shuyixiong	6776caaad1	[TRTLLM-8507][fix] Fix ray resource cleanup and error handling in LoRA test (#8175 ) Signed-off-by: shuyix <219646547+shuyixiong@users.noreply.github.com>	2025-10-14 23:46:30 +08:00
Fanrong Li	0d20a8fd61	[TRTLLM-8536][feat] Add the sparse attention framework and one use case--RocketKV support (#8086 ) Signed-off-by: Fanrong Li <23290157+lfr-0531@users.noreply.github.com> Signed-off-by: yuhangh <58161490+heyuhhh@users.noreply.github.com> Co-authored-by: yuhangh <58161490+heyuhhh@users.noreply.github.com>	2025-10-14 08:23:16 -07:00
Yan Chunwei	86be06bda4	[None][ci] waive several rpc tests (#8349 ) Signed-off-by: Yan Chunwei <328693+Superjomn@users.noreply.github.com>	2025-10-14 03:12:49 -07:00
William Zhang	72d65d079a	[https://nvbugs/5542878 ][fix] Unwaive test (#8027 ) Signed-off-by: William Zhang <133824995+2ez4bz@users.noreply.github.com>	2025-10-14 07:58:07 +02:00
xinhe-nv	371fcb0338	[TRTLLM-8366][feat] add kimi multi nodes case (#8025 ) Signed-off-by: Xin He (SW-GPU) <200704525+xinhe-nv@users.noreply.github.com>	2025-10-13 21:36:03 -07:00
Yuxian Qiu	3450fe9944	[None][fix] Fix dummy load format for key models. (#7993 ) Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com>	2025-10-14 11:18:39 +08:00
Lucas Liebenwein	22aa4ac08c	[None][feat] AutoDeploy: VLMs with subgraphs + cudagraph/compile (#8203 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-10-13 17:34:09 -07:00
Zheyu Fu	bac665e650	[TRTLLM-7412][feat] Turn off spec decode when the rolling average acceptance length drops below threshold. (#7283 ) Signed-off-by: Zheyu Fu <zheyuf@NVIDIA.com>	2025-10-13 15:51:14 -07:00
Robin Kobus	db8c63b9b1	[TRTLLM-4517] [feat] Additional model outputs (#7206 ) Signed-off-by: Robin Kobus <19427718+Funatiq@users.noreply.github.com>	2025-10-13 15:33:18 +02:00
amitz-nv	bbae7a05f0	[https://nvbugs/5521949 ][fix] Replace test_codellama_fp8_with_bf16_lora with test_llama_3_1_8b_fp8_with_bf16_lora (#8199 ) Signed-off-by: Amit Zuker <203509407+amitz-nv@users.noreply.github.com>	2025-10-13 06:01:55 -07:00
Po-Han Huang (NVIDIA)	6fc6f70a68	[https://nvbugs/5441729 ][test] Fix test_modeling_llama_min_latency.py failures (#7478 ) Signed-off-by: Po-Han Huang <pohanh@nvidia.com>	2025-10-13 15:35:02 +08:00
xinhe-nv	9fe63dd8db	[None][chore] Add failed cases into waives.txt (#8290 ) Signed-off-by: xinhe-nv <200704525+xinhe-nv@users.noreply.github.com>	2025-10-13 00:07:00 -07:00
Leslie Fang	8d1b068b1a	[TRTLLM-8477][chore] Replace KvCacheConfigCpp with KvCacheConfig inside PyExecutor (#8259 ) Signed-off-by: leslie-fang25 <leslief@nvidia.com>	2025-10-13 14:55:36 +08:00
xinhe-nv	72fcff1044	[None][fix] add timeout for llama4 (#8254 ) Signed-off-by: Xin He (SW-GPU) <200704525+xinhe-nv@users.noreply.github.com>	2025-10-12 21:04:20 -07:00
Guoming Zhang	989c25fcba	[None][doc] Add qwen3-next doc into deployment guid and test case into L0. (#8288 ) Signed-off-by: nv-guomingz <137257613+nv-guomingz@users.noreply.github.com> Co-authored-by: Faradawn Yang <faradawny@gmail.com> Co-authored-by: Robin Kobus <19427718+Funatiq@users.noreply.github.com>	2025-10-13 10:25:45 +08:00
amitz-nv	fac47e2826	[https://nvbugs/5510879 ][fix] Fix pytorch & TRT-python flows fused LoRA adapter modules weight split with TP>1 (#8063 ) Signed-off-by: Amit Zuker <203509407+amitz-nv@users.noreply.github.com>	2025-10-12 12:29:52 -07:00
Eran Geva	a1ed03fe8a	[None][fix] AD test_trtllm_bench to use small model config and skip loading weights (#8149 ) Signed-off-by: Eran Geva <19514940+MrGeva@users.noreply.github.com>	2025-10-12 18:30:20 +03:00
Emma Qiao	fdbeea51d3	[None][infra] Skip failed cases for main branch (#8293 ) Signed-off-by: qqiao <qqiao@nvidia.com>	2025-10-12 08:04:09 -07:00
kris1025	a7ea544dbe	[TRTLLM-7384][feat] enable rejection sampling for CDL (#7731 ) Signed-off-by: linquanh <linquanh@nvidia.com>	2025-10-12 20:38:48 +08:00
brb-nv	56a539cd37	[None][chore] Waive failing pre-merge test on main (#8282 ) Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>	2025-10-10 23:52:05 -07:00
Yilin Fan	2695d70d42	[None][feat] Add request timing breakdown option in benchmark_serving (#8128 ) Signed-off-by: nv-yilinf <206948969+nv-yilinf@users.noreply.github.com>	2025-10-10 09:24:54 -07:00
xinhe-nv	2655995a09	[None][fix] add gc for test fixture (#8220 ) Signed-off-by: Xin He (SW-GPU) <200704525+xinhe-nv@users.noreply.github.com>	2025-10-10 02:50:25 -07:00
bhsueh_NV	d3059dbd8a	[https://nvbugs/5547416 ][fix] unwaive no_cache test (#8213 ) Signed-off-by: bhsueh <11360707+byshiue@users.noreply.github.com>	2025-10-10 01:50:13 -07:00
xinhe-nv	b555f1ff98	[None][chore] Add failed cases into waives.txt (#8229 ) Signed-off-by: Xin He (SW-GPU) <200704525+xinhe-nv@users.noreply.github.com>	2025-10-09 23:45:28 -07:00
xinhe-nv	e8c9bae37e	[None][chore] Remove closed bugs (#8151 ) Signed-off-by: Xin He (SW-GPU) <200704525+xinhe-nv@users.noreply.github.com>	2025-10-10 16:39:40 +11:00
Pengbo Wang	7da4b05289	[https://nvbugs/5501820 ][fix] Add requirements for numba-cuda version to WAR mem corruption (#7992 ) Signed-off-by: Pengbo Wang <221450789+pengbowang-nv@users.noreply.github.com>	2025-10-10 10:18:27 +08:00
Emma Qiao	ccd949ea5b	[None][infra] Waive failed tests on main 10/09 (#8230 ) Signed-off-by: qqiao <qqiao@nvidia.com>	2025-10-09 22:46:07 +08:00
amitz-nv	d560054e1b	[None][chore] Restore asserts in pytorch flow LoRA tests (#8227 ) Signed-off-by: Amit Zuker <203509407+amitz-nv@users.noreply.github.com>	2025-10-09 17:10:38 +03:00
bhsueh_NV	27677a36f5	[https://nvbugs/5516666 ][fix] unwaive some Qwen3 CI tests (#8130 ) Signed-off-by: bhsueh <11360707+byshiue@users.noreply.github.com>	2025-10-09 09:44:58 +08:00
Lizhi Zhou	fdf29ab8fa	[TRTLLM-7846][feat] Http disagg-cluster management implemention (#7869 ) Signed-off-by: Lizhi Zhou <1432185+reasonsolo@users.noreply.github.com>	2025-10-09 09:44:01 +08:00
QI JUN	6884d06aed	[None][ci] move some llama4 test cases to pre merge (#8189 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>	2025-10-08 18:34:08 -07:00
Liao Lanyu	ed8e00ad4a	[https://nvbugs/5522746 ][fix] unwaive tests caused by node issues after rebooting (#8193 ) Signed-off-by: Lanyu Liao <lancelly@users.noreply.github.com> Co-authored-by: Lanyu Liao <lancelly@users.noreply.github.com>	2025-10-09 08:45:56 +08:00
Mike Iovine	c88913dc03	[https://nvbugs/5541545 ][fix] Remove test_llama4 (#8031 ) Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com>	2025-10-08 15:20:15 -07:00
brb-nv	80517b7812	[None][chore] Waive some tests failing on main post merge (#8186 ) Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>	2025-10-08 06:52:30 -07:00
mpikulski	8298e93bd8	[TRTLLM-8414][chore] BREAKING CHANGE: refine sampling strategy selection (#8132 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-10-08 15:46:50 +02:00
xxi	e98616512f	[https://nvbugs/5550283 ][fix] update test case to the latest MoE API (#8165 )	2025-10-07 22:54:34 -07:00
Liao Lanyu	d57b8f0951	[https://nvbugs/5455140 ][fix] unwaive tests related to GB200 OOM (#8159 ) Signed-off-by: Lanyu Liao <lancelly@users.noreply.github.com> Co-authored-by: Lanyu Liao <lancelly@users.noreply.github.com>	2025-10-08 13:14:12 +08:00
ruodil	971610e3ff	[None][test] add test-model-suites option in integration conftest.py (#8016 ) Signed-off-by: Ruodi Lu <ruodil@users.noreply.github.com> Co-authored-by: Ruodi Lu <ruodil@users.noreply.github.com>	2025-10-08 10:38:31 +08:00
Mike Iovine	7facac077b	[None][fix] Fix MTP illegal memory access (#8161 ) Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com>	2025-10-07 14:02:55 -04:00
Emma Qiao	ca9da1f1c2	[None][infra] Skip failed cases for main (#8176 ) Signed-off-by: qqiao <qqiao@nvidia.com>	2025-10-07 06:37:51 -07:00
xiweny	9298f1bdcc	[None] [test] Add B300 cases to CI (#8056 ) Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com>	2025-10-06 19:23:31 -07:00
Faraz	27a5091fcb	[None][feat] GPT-OSS Sm120/Sm121 Support (#7937 ) Signed-off-by: Perkz Zheng <67892460+PerkzZheng@users.noreply.github.com> Signed-off-by: list <58580514+farazkh80@users.noreply.github.com> Signed-off-by: Vincent Huang <vincenth@nvidia.com> Co-authored-by: Perkz Zheng <67892460+PerkzZheng@users.noreply.github.com> Co-authored-by: Vincent Huang <vincenth@nvidia.com>	2025-10-06 16:59:06 -04:00
Izzy Putterman	f2657c1ae9	[None][fix] Eagle: Attention DP (#7939 ) Signed-off-by: Izzy Putterman <iputterman@nvidia.com>	2025-10-06 16:52:35 -04:00

1 2 3 4 5 ...

1695 Commits