TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-02-05 02:31:33 +08:00

Author	SHA1	Message	Date
shuyixiong	fd2af8d58a	[TRTLLM-9771][feat] Support partial update weight for fp8 (#10456 ) Signed-off-by: Shuyi Xiong <219646547+shuyixiong@users.noreply.github.com> Signed-off-by: shuyixiong <219646547+shuyixiong@users.noreply.github.com>	2026-01-22 14:46:05 +08:00
Wanli Jiang	ff0775408d	[None][fix] Fix waived tests for Nemotron-h models (#10758 ) Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>	2026-01-22 14:17:50 +08:00
Enwei Zhu	be4a431ffd	[TRTLLM-10154][feat] Enable guided decoding with reasoning parsers (#10890 ) Signed-off-by: Enwei Zhu <21126786+syuoni@users.noreply.github.com>	2026-01-22 14:14:28 +08:00
Taylor Yeonbok Lee	895bb94b3d	[#8241 ][feat] Support model_kwargs for pytorch backend (#10351 ) Signed-off-by: Taylor Yeonbok Lee <249374542+taylor-yb-lee@users.noreply.github.com>	2026-01-21 20:51:38 -08:00
Yechan Kim	70caa779a4	[None][feat] K-EXAONE MTP support (#10796 ) Signed-off-by: yechank <161688079+yechank-nvidia@users.noreply.github.com>	2026-01-22 13:43:00 +09:00
JennyLiu	415739711f	[None][chore] Add DGX-Spark VLM accuracy and perf spec dec cases (#10804 ) Signed-off-by: Jenny Liu <JennyLiu-nv+JennyLiu@users.noreply.github.com> Signed-off-by: JennyLiu <141791095+JennyLiu-nv@users.noreply.github.com> Co-authored-by: Jenny Liu <JennyLiu-nv+JennyLiu@users.noreply.github.com>	2026-01-22 12:38:17 +08:00
Lizhi Zhou	f3a41c8d94	[TRTLLM-10059][feat] Use global unique id as disagg request id (#10187 ) Signed-off-by: Lizhi Zhou <1432185+reasonsolo@users.noreply.github.com>	2026-01-21 22:52:34 -05:00
Daniil	0434db5bf7	[None][feat] GLM-4.5-Air support (#10653 ) Signed-off-by: Daniil Kulko <kulkodaniil@gmail.com>	2026-01-22 11:42:09 +08:00
TensorRT LLM	bd56b4e1e3	[None][infra] Check in most recent lock file from nightly pipeline Signed-off-by: TensorRT LLM <90828364+tensorrt-cicd@users.noreply.github.com>	2026-01-22 03:24:57 +00:00
Yuxian Qiu	c2a9e66dff	[https://nvbugs/5784543 ][chore] unwaive test. (#10835 ) Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com>	2026-01-22 11:17:28 +08:00
dongxuy04	635cbf01ba	[https://nvbugs/5816267 ][fix] Remove weight tensor holder to release memory earlier (#10876 ) Signed-off-by: Dongxu Yang <78518666+dongxuy04@users.noreply.github.com>	2026-01-21 16:42:52 -08:00
yuanjingx87	5450485bec	[None][infra] Fix sonarQube job hang by create jenkins homd folder if not exist (#10830 ) Signed-off-by: Yuanjing Xue <197832395+yuanjingx87@users.noreply.github.com>	2026-01-21 11:45:19 -08:00
Guiju Zhang	8cf8fbbe16	[TRTLLM-10325][feat] Refactor speculative decoding workers (#10768 ) Signed-off-by: Guiju Zhang <7135567+cascade812@users.noreply.github.com>	2026-01-21 13:05:29 -05:00
kris1025	f91ea37a13	[None][chore] unwaive qwen3 235B accuracy test (#10493 ) Signed-off-by: linquanh <linquanh@nvidia.com>	2026-01-21 17:52:04 +08:00
Yukun He	bf7303c7f1	[https://nvbugs/5636916 ][fix] Cherry-pick #10654 : Fix accuracy issue of TWO-SHOT AllReduce kernel (#10841 ) Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>	2026-01-21 17:25:40 +08:00
Emma Qiao	165dd360b9	[None][infra] Waive failed cases for main branch on 01/21 (#10882 ) Signed-off-by: qqiao <qqiao@nvidia.com>	2026-01-21 04:24:05 -05:00
xxi	9feebb3a27	[None][chore] switch to ConfigurableMoE as the default path (#10792 ) Signed-off-by: xxi <xxi@nvidia.com>	2026-01-21 15:57:38 +08:00
Yukun He	a4152c80f6	[https://nvbugs/5814253 ][fix] unwaive test_autotuner_distributed_strategy tests (#10793 ) Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>	2026-01-21 15:37:11 +08:00
HuiGao-NV	1592dfab6d	[https://nvbugs/5740377 ][fix] Lock resource to fix potential access to released data (#10827 ) Signed-off-by: Hui Gao <huig@nvidia.com>	2026-01-21 14:17:29 +08:00
Yukun He	d60d6ff6fd	[None][fix] Cherry-pick #10715 : Disable short profile for tunable ops with MERGE strategy (#10844 ) Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>	2026-01-21 13:53:04 +08:00
Xianjie Qiao	87073d1ce4	[None][fix] Fix copy start_logs in disagg slurm scripts (#10840 ) Signed-off-by: Xianjie <5410381+qiaoxj07@users.noreply.github.com>	2026-01-21 13:31:25 +08:00
Yibin Li	9116dfbacd	[https://nvbugs/5775021 ] [fix] Replace pickle.load with restricted Unpickler (#10622 ) Signed-off-by: Yibin Li <109242046+yibinl-nvidia@users.noreply.github.com>	2026-01-21 11:42:54 +08:00
TensorRT LLM	ffd2ed51dd	[None][infra] Check in most recent lock file from nightly pipeline Signed-off-by: TensorRT LLM <90828364+tensorrt-cicd@users.noreply.github.com>	2026-01-21 03:14:33 +00:00
Yanchao Lu	ccf4d79c6c	[None][chore] Revert NVIDIA/TensorRT-LLM#10847 (#10869 )	2026-01-21 11:08:40 +08:00
shuyixiong	c381790d15	[https://nvbugs/5670458 ][chore] Unwaive reward model test (#10831 ) Signed-off-by: shuyix <219646547+shuyixiong@users.noreply.github.com>	2026-01-21 10:34:01 +08:00
Daniel Stokes	2f3b2a3172	[None][fix] Add a timeout in MNNVL throughput to prevent hangs if one rank crashes (#9532 ) Signed-off-by: djns99 <40156487+djns99@users.noreply.github.com> Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com> Co-authored-by: Bo Li <22713281+bobboli@users.noreply.github.com>	2026-01-21 10:14:39 +08:00
Yan Chunwei	3c39b1faa9	[https://nvbugs/5759698 ][fix] unwaive test_base_worker (#10669 ) Signed-off-by: Yan Chunwei <328693+Superjomn@users.noreply.github.com>	2026-01-20 21:14:03 -05:00
Zheng Duan	26c23cf99f	[https://nvbugs/5760737 ][test] only skip mooncake+indexerkcache test (#10266 ) Signed-off-by: zhengd-nv <200704041+zhengd-nv@users.noreply.github.com>	2026-01-21 09:48:39 +08:00
Simeng Liu	3c8ed19440	[https://nvbugs/5670108 ][fix] Fix overlap scheduler race condition in… (#10610 ) Signed-off-by: SimengLiu-nv <simengl@nvidia.com>	2026-01-20 10:56:56 -08:00
TensorRT LLM	c6163e2b70	[None][infra] Check in most recent lock file from nightly pipeline Signed-off-by: TensorRT LLM <90828364+tensorrt-cicd@users.noreply.github.com>	2026-01-20 18:36:19 +00:00
Izzy Putterman	864b61cadd	[None][feat] Speculative One Model: FlashInfer sampling (#10284 ) Signed-off-by: Izzy Putterman <iputterman@nvidia.com>	2026-01-20 12:56:43 -05:00
Lucas Liebenwein	66b239a9a9	[None][fix] fix duplicate entry in waives.txt (#10853 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2026-01-20 19:48:01 +02:00
jthomson04	2db3d7eeba	[None][chore] Async Transfer Manager (#9891 ) Signed-off-by: jthomson04 <jwillthomson19@gmail.com>	2026-01-20 12:12:47 -05:00
Gal Hubara-Agam	e61c942d1f	[#10707 ][fix] AutoDeploy: Super accuracy test fixes (#10717 ) Signed-off-by: Gal Hubara Agam <96368689+galagam@users.noreply.github.com> Signed-off-by: Gal Hubara-Agam <96368689+galagam@users.noreply.github.com>	2026-01-20 18:16:13 +02:00
Yanchao Lu	ae8f74b620	[None][chore] Reduce tedious logs (#10847 ) Signed-off-by: Yanchao Lu <yanchaol@nvidia.com>	2026-01-20 22:56:24 +08:00
Emma Qiao	3a894951e7	[None][infra] Waive failed cases for main branch on 01/20 (#10829 ) Signed-off-by: qqiao <qqiao@nvidia.com>	2026-01-20 17:58:58 +08:00
Bo Deng	338b29d5ae	[None][infra] trigger multi-gpu tests when install_nixl/ucx.sh is mod… (#10624 ) Signed-off-by: Bo Deng <deemod@nvidia.com>	2026-01-20 17:55:32 +08:00
Yuxian Qiu	c8a200486d	[https://nvbugs/5701445 ][chore] unwaive test. (#10806 ) Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com>	2026-01-20 16:30:32 +08:00
Grzegorz Kwasniewski	eb326073d8	[TRTLLM-10785][feat] Fix sharding dashboard errors (#10786 ) Signed-off-by: greg-kwasniewski1 <213329731+greg-kwasniewski1@users.noreply.github.com>	2026-01-20 09:25:36 +01:00
Yi Zhang	58311b2345	[None][fix] Remove unused params in attn (#10652 ) Signed-off-by: yizhang-nv <187001205+yizhang-nv@users.noreply.github.com>	2026-01-20 03:08:59 -05:00
xinhe-nv	47e0ec2527	[None][test] Update sanity test list (#10825 ) Signed-off-by: Xin He (SW-GPU) <200704525+xinhe-nv@users.noreply.github.com>	2026-01-20 02:11:42 -05:00
Yiqing Yan	99e8cb0999	[None][fix] Fix vulnerability urllib3 and nbconvert (#10551 ) Signed-off-by: Yiqing Yan <yiqingy@nvidia.com>	2026-01-20 14:51:36 +08:00
xinhe-nv	fc467d06c3	[TRTLLM-8638][fix] Add failed cases into waives.txt (#10787 ) Signed-off-by: xinhe-nv <200704525+xinhe-nv@users.noreply.github.com> Signed-off-by: Xin He (SW-GPU) <200704525+xinhe-nv@users.noreply.github.com>	2026-01-20 00:48:19 -05:00
benzh-2025	4c8468c5d3	[None][fix] default disable gemm+allreduce fusion (#10656 )	2026-01-20 12:31:17 +08:00
xinhe-nv	26bc16842e	[None][chore] Add failed cases into waives.txt (#10776 ) Signed-off-by: Jie Li <lijie@nvidia.com> Co-authored-by: Jie Li <lijie@nvidia.com>	2026-01-19 22:45:40 -05:00
TensorRT LLM	44c5af88dc	[None][infra] Check in most recent lock file from nightly pipeline Signed-off-by: TensorRT LLM <90828364+tensorrt-cicd@users.noreply.github.com>	2026-01-20 03:15:53 +00:00
Bo Li	f3a985ce27	[TRTLLM-10296][fix] Fix the potential misaligned access due to vectorized ld/st instructions in NVLinkOneSided A2A. (#10539 ) Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>	2026-01-20 11:08:04 +08:00
Liao Lanyu	dbb858ae0c	[TRTLLM-10029][scheduler] Re-implement MicroBatchScheduler and CapacityScheduler in Python (#10273 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com> Signed-off-by: Lanyu Liao <lancelly@users.noreply.github.com> Signed-off-by: Lance Liao <108499334+lancelly@users.noreply.github.com> Co-authored-by: junq <22017000+QiJune@users.noreply.github.com> Co-authored-by: Lanyu Liao <lancelly@users.noreply.github.com>	2026-01-20 10:31:13 +08:00
Lizhi Zhou	c6320d924d	[https://nvbugs/5776445 ][chore] unwaive test (#10667 ) Signed-off-by: Lizhi Zhou <1432185+reasonsolo@users.noreply.github.com>	2026-01-19 21:22:47 -05:00
Zhenhuan Chen	066fa4cd93	[None][chore] update config.yaml of slurm scripts to align with submit.py change (#10802 ) Signed-off-by: Zhenhuan Chen <zhenhuanc@nvidia.com>	2026-01-19 14:46:23 -05:00

1 2 3 4 5 ...

4766 Commits