TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

Author	SHA1	Message	Date
dongfengy	f5575a9146	[https://nvbugs/5474119 ][fix] Cherry-pick https://github.com/NVIDIA/TensorRT-LLM/pull/8809 (#8847 ) Signed-off-by: Dongfeng Yu <dongfengy@nvidia.com>	2025-11-02 16:44:06 -08:00
Barry Kang	f22a87f296	[https://nvbugs/5325296 ][fix] Enable relaxed acceptance test on Blackwell (#8709 ) Signed-off-by: Barry Kang <43644113+Barry-Delaney@users.noreply.github.com>	2025-10-31 15:02:06 -07:00
Lucas Liebenwein	752cc3a8cb	[https://nvbugs/5606166 ][fix] AutoDeploy: use tuples for cudagraph shape lookup (#8772 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-10-31 13:59:48 +01:00
Zhanrui Sun	d2071d7ed7	[None][infra] Remove invaild waived tests which not in release branch (#8841 ) Signed-off-by: ZhanruiSunCh <184402041+ZhanruiSunCh@users.noreply.github.com>	2025-10-31 03:02:34 -07:00
Emma Qiao	421d48f402	[None][infra] Skip failed tests for release branch (#8833 ) Signed-off-by: qqiao <qqiao@nvidia.com>	2025-10-31 15:04:54 +08:00
Jin Li	28673f3e9c	[https://nvbugs/5488118 ][fix] Unwaive passed tests (#8758 ) Signed-off-by: Jin Li <59594262+liji-nv@users.noreply.github.com>	2025-10-31 10:46:44 +08:00
Emma Qiao	9ee0075921	[TRTLLM-8971][infra] Cherry-pick for Update gpu key for B300/GB300 (#8724 ) (#8796 ) Signed-off-by: qqiao <qqiao@nvidia.com>	2025-10-30 06:12:16 -07:00
Dom Brown	9410ce3bea	[https://nvbugs/5575841 ] [test] Move test_moe.py to serial tests to improve stability + unwaive FP4 MoE torch unit tests (#8422 ) Signed-off-by: Dom Brown <3886319+DomBrown@users.noreply.github.com>	2025-10-30 13:57:56 +01:00
Emma Qiao	ec510ad72a	[None][infra] Waive failed tests for release branch (#8760 ) Signed-off-by: qqiao <qqiao@nvidia.com>	2025-10-29 06:32:30 -07:00
xiweny	f49f42db59	[https://nvbugs/5601203 ] [fix]Restrict fp8 blockscale moe case (#8583 ) Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com>	2025-10-29 10:47:32 +08:00
Eran Geva	db3c373d3a	[https://nvbugs/5572320 ][fix] Ported test_ad_trtllm_bench.py from main (#8671 ) Signed-off-by: Eran Geva <19514940+MrGeva@users.noreply.github.com>	2025-10-28 09:41:32 +02:00
Yukun He	e04354bc09	[https://nvbugs/5608489 ][fix] Fix output unpack issues for Llama3/4 NVFP4 models. (#8679 ) Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>	2025-10-28 14:21:47 +08:00
Shiyu Li	28c9a51c06	[https://nvbugs/5597647 ][fix] Fix MNNVL Allreduce accuracy issue on Hopper (#8612 ) Signed-off-by: Shiyu Li <shili@nvidia.com>	2025-10-26 23:06:45 -07:00
Emma Qiao	b05555faeb	[None][infra] Waive failed tests for release 10/24 (#8656 ) Signed-off-by: qqiao <qqiao@nvidia.com> Signed-off-by: Emma Qiao <qqiao@nvidia.com>	2025-10-24 21:53:35 +08:00
Ivy Zhang	1859b55d22	[None][test] Clean cache for certain easily hang cases (#8619 ) Signed-off-by: Ivy Zhang <25222398+crazydemo@users.noreply.github.com> Co-authored-by: Larry Xu <197874197+LarryXFly@users.noreply.github.com>	2025-10-24 08:17:32 -04:00
Jie Li	7749ec406b	[https://nvbugs/5587456 ][fix] Remove multimodal test cases using TRT backend (#8611 ) Signed-off-by: Jie Li <lijie@nvidia.com>	2025-10-24 18:04:43 +08:00
Jie Li	4b52054bdd	[https://nvbugs/5541145 ][fix] Remove DeepSeekR1 test case from H20 to prevent OOM (#8610 ) Signed-off-by: Jie Li <lijie@nvidia.com>	2025-10-24 05:20:40 -04:00
Leslie Fang	d9d898e8b7	[https://nvbugs/5608461 ][fix] exclude InductorSubproc from thread leak check (#8624 ) Signed-off-by: leslie-fang25 <leslief@nvidia.com>	2025-10-24 13:08:42 +08:00
Zheyu Fu	d2c976ceac	[https://nvbugs/5576192 ][fix] Unwaive the test for test_weight_only_quant_gemm. (#8546 ) Signed-off-by: Zheyu Fu <zheyuf@NVIDIA.com>	2025-10-23 15:46:09 -07:00
Lizhi Zhou	686298d2d5	[https://nvbugs/5575902 ][fix] set max_batch_size=1 to stabilize accuracy test result (#8609 ) Signed-off-by: Lizhi Zhou <1432185+reasonsolo@users.noreply.github.com>	2025-10-23 07:28:29 -07:00
Ivy Zhang	5d27034295	[TRTLLM-8785][fix] create output_dir before test begin (cherry-pick #8518 ) (#8575 ) Signed-off-by: Ivy Zhang <25222398+crazydemo@users.noreply.github.com>	2025-10-23 04:41:54 -04:00
Chang Liu	e5b6d335eb	[https://nvbugs/5568961 ][fix] Fix a merge conflict (cherrypick from PR 8365) (#8553 ) Signed-off-by: Jin Li <59594262+liji-nv@users.noreply.github.com> Co-authored-by: Jin Li <59594262+liji-nv@users.noreply.github.com>	2025-10-23 14:05:16 +08:00
Lizhi Zhou	3f82cdbdad	[https://nvbugs/5582277 ][fix] rework DisaggPPTerminationHandler to fix hang issue (#8519 ) Signed-off-by: Lizhi Zhou <1432185+reasonsolo@users.noreply.github.com>	2025-10-23 09:43:59 +08:00
Pengyun Lin	e86d6db9ec	[https://nvbugs/5575829 ][fix] Unwaive gpt-oss test (#8576 ) Signed-off-by: Pengyun Lin <81065165+LinPoly@users.noreply.github.com>	2025-10-22 07:31:56 -04:00
Emma Qiao	09349ccbfe	[None][infra] Waive failed tests for release 10/22 (#8574 ) Signed-off-by: qqiao <qqiao@nvidia.com>	2025-10-22 04:41:00 -04:00
Bo Deng	9e30f14da8	[https://nvbugs/5565549 ][fix] unwaive test_disaggregated_spec_dec_bat… (#8500 ) Signed-off-by: Bo Deng <deemod@nvidia.com>	2025-10-22 14:59:59 +08:00
Guoming Zhang	a519c2c43c	[https://nvbugs/5504095 ][fix] Unwaive test_user_specify_workspace case. (#8316 ) Signed-off-by: nv-guomingz <137257613+nv-guomingz@users.noreply.github.com>	2025-10-22 09:31:24 +08:00
Simeng Liu	1375b9f074	[https://nvbugs/5515753 ][ci] Add NCCL_DEBUG=INFO flag to collect more info with CI failure. (#8440 ) Signed-off-by: Simeng Liu <simengl@nvidia.com>	2025-10-21 18:12:05 -07:00
JunyiXu-nv	0acdecb2c3	[https://nvbugs/5569713 ][fix] Disable fp8 deep gemm for EXAONE-4.0-32B-FP8 (#8429 ) Signed-off-by: Junyi Xu <219237550+JunyiXu-nv@users.noreply.github.com>	2025-10-21 12:37:56 -04:00
mpikulski	f256eb9063	[TRTLLM-8650][fix] beam search request validation (#8433 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-10-21 10:50:27 +02:00
Emma Qiao	2b0a10e4d5	[None][infra] Waive tests for release 1021 (#8522 ) Signed-off-by: qqiao <qqiao@nvidia.com>	2025-10-21 03:21:00 -04:00
Pengbo Wang	8ce2dc5cb7	[https://nvbugs/5501820 ][fix] Add requirements for numba-cuda version to WAR mem corruption (#7992 ) (#8414 ) Signed-off-by: Pengbo Wang <221450789+pengbowang-nv@users.noreply.github.com>	2025-10-20 09:01:08 +02:00
bhsueh_NV	14d0f5d683	[https://nvbugs/5516666 ][fix] cherry-pick PR 8130 to unwaive the Qwen3 CI (#8444 ) Signed-off-by: bhsueh <11360707+byshiue@users.noreply.github.com>	2025-10-19 23:14:10 -04:00
Ivy Zhang	f904348cd6	[TRTLLM-8580][test] save runtime report periodically (#8312 ) (#8455 ) Signed-off-by: Ivy Zhang <25222398+crazydemo@users.noreply.github.com>	2025-10-20 10:54:24 +08:00
xiweny	af2450c266	[https://nvbugs/5565565 ] [fix] Remove waiver (#8450 ) Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com>	2025-10-17 01:13:01 -07:00
Yukun He	437a3fc642	[None][chore] Remove duplicate log outputs in test_perf.py (#8418 ) Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>	2025-10-17 14:11:32 +08:00
Yan Chunwei	995b93bc38	[https://nvbugs/5437384 ][test] fix trtllm-llmapi-launch multi tests with single launch (#8397 ) Signed-off-by: Yan Chunwei <328693+Superjomn@users.noreply.github.com> Signed-off-by: Superjomn <328693+Superjomn@users.noreply.github.com>	2025-10-16 21:14:43 -07:00
ruodil	20c2de4924	[None][test] cherry-pick: add test-model-suites in integration conftest.py (#8388 ) Signed-off-by: Ruodi Lu <ruodil@users.noreply.github.com> Co-authored-by: Ruodi Lu <ruodil@users.noreply.github.com> Co-authored-by: Larry <197874197+LarryXFly@users.noreply.github.com>	2025-10-15 23:26:32 -07:00
Yukun He	fd4311e6a3	[TRTLLM-8129][feat] Allreduce tuning and benchmark script revising (#7870 ) Because we have encountered some perf regression due to using a one-shot kernel instead of NCCL on A100/H100, it will be beneficial if we can have a solid benchmarking of allreduce Op and analyze the data collected from it. Implemented new AllreduceOp heuristic: - Added Linear programming-based heuristic implementation. - Added LUT-based heuristic implementation and corresponding code generation script. AllreduceOp minor fixing: - Fixed a minor issue in AllreduceOp, that the strategy can not be overridden when ONESHOT or TWOSHOT is set. - Fixed a minor TWOSHOT kernel perf issue. - Cleaned up Dispatching code in AllReduceOp. This PR will fix the perf gaps reported in: https://nvbugspro.nvidia.com/bug/5517023 For Deepseek-R1, it shows a performance gain of about 3-4% in concurrency levels of 256 and 512. Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>	2025-10-16 14:15:25 +08:00
Patrice Castonguay	7862372ee2	[https://nvbugs/5552889 ][fix] fix: Prevent empty batch when using attention DP with disagg (#8372 ) Signed-off-by: Patrice Castonguay <55748270+pcastonguay@users.noreply.github.com>	2025-10-16 09:11:04 +08:00
amitz-nv	27c6c8466b	[https://nvbugs/5510879 ][fix] Fix pytorch & TRT-python flows fused LoRA adapter modules weight split with TP>1 (#8313 ) Signed-off-by: Amit Zuker <203509407+amitz-nv@users.noreply.github.com>	2025-10-15 08:24:02 -07:00
amitz-nv	e5476a6b2a	[https://nvbugs/5521949 ][fix] Update FP8 model with BF16 LoRA test, fix test_bielik_11b_v2_2_instruct_multi_lora (#8324 ) Signed-off-by: Amit Zuker <203509407+amitz-nv@users.noreply.github.com>	2025-10-15 05:48:38 -07:00
Ivy Zhang	4751bdbcb6	[None][chore] Update nim test list (#8356 ) Signed-off-by: Ivy Zhang <25222398+crazydemo@users.noreply.github.com> Co-authored-by: Larry <197874197+LarryXFly@users.noreply.github.com>	2025-10-15 02:04:20 -07:00
Emma Qiao	988f93790f	[None][infra] Waive failed tests in release post-merge 10/15 (#8386 ) Signed-off-by: qqiao <qqiao@nvidia.com>	2025-10-15 16:06:08 +08:00
Stanley Sun	cce97e6e15	[TRTLLM-8113][test] Add pytorch workflow e2e tests with pp enabled (#8357 ) Signed-off-by: Stanley Sun <stsun@nvidia.com>	2025-10-15 15:09:21 +08:00
xiweny	d5b79268e7	[https://nvbugs/5565565 ] [fix] fp8 wideep support sm103 (#8228 ) Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com>	2025-10-15 10:17:08 +08:00
Yiqing Yan	7b5ba7ca66	[https://nvbugs/5565541 ][fix] Add timeout threshold for H100 FHMA test (#8354 ) Signed-off-by: Yiqing Yan <yiqingy@nvidia.com>	2025-10-14 01:23:08 -07:00
bhsueh_NV	66aa88739b	[https://nvbugs/5574556 ][fix] fix bug of Qwen3_235B_A22B::test_fp8 CI (#8351 ) Signed-off-by: bhsueh <11360707+byshiue@users.noreply.github.com>	2025-10-14 15:26:15 +08:00
Ziyi Xiong	9ecc6db5b4	[https://nvbugs/5537878 ][fix] Reserve an extra slot for padded batch … (#8231 ) Signed-off-by: ziyixiong-nv <219238287+ziyixiong-nv@users.noreply.github.com>	2025-10-13 23:34:22 -07:00
Lizhi Zhou	553ff3402a	[https://nvbugs/5550671 ][fix] fix disagg-serving multinodes test failure (#8307 ) Signed-off-by: Lizhi Zhou <1432185+reasonsolo@users.noreply.github.com>	2025-10-14 08:01:00 +02:00

1 2 3 4 5 ...

1656 Commits