TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

Author	SHA1	Message	Date
Zero Zeng	4545700fcf	[None][chore] Move submit.sh to python and use yaml configuration (#8003 ) Signed-off-by: Zero Zeng <38289304+zerollzeng@users.noreply.github.com>	2025-10-20 22:36:50 -04:00
mpikulski	87eb5086fb	[None][fix] restore list[list[list[int]]] in add_token (#8502 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-10-20 22:34:57 -04:00
Yechan Kim	85d5aa7763	[None][feat] Support kv_cahce_reuse for HyperCLOVAX-Vision model (#7789 ) Signed-off-by: yechank <161688079+yechank-nvidia@users.noreply.github.com>	2025-10-21 11:11:24 +09:00
Ruoqian Guo	984d4fe0fe	[None][feat] Update 3rdparty/DeepGEMM to latest commit (#8488 ) Signed-off-by: Ruoqian Guo <22525902+ruoqianguo@users.noreply.github.com> Co-authored-by: Barry Kang <43644113+Barry-Delaney@users.noreply.github.com>	2025-10-21 06:56:50 +08:00
Suyog Gupta	7050b1ea49	[#8272 ][feat] Enable chunked prefill for SSMs in AutoDeploy (#8477 ) Signed-off-by: Suyog Gupta <41447211+suyoggupta@users.noreply.github.com>	2025-10-20 15:31:52 -07:00
Venky	3e681e2a80	[None] [chore] Add architecture-specific ATTRIBUTIONS files (#8468 ) Signed-off-by: Venky Ganesh <23023424+venkywonka@users.noreply.github.com>	2025-10-20 16:29:15 -04:00
Lucas Liebenwein	55c468b218	[#8461 ][feat] AutoDeploy: trtllm-serve bug fix + unit test (#8462 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-10-20 16:06:39 -04:00
dongfengy	9b289d5230	[https://nvbugs/5568676 ][fix] Remove test waive (#8437 ) Signed-off-by: Dongfeng Yu <dongfengy@nvidia.com>	2025-10-20 12:03:50 -07:00
yuanjingx87	1e3e1474c6	[TRTLLM-6055][infra] Slurm Test refactor (#7176 ) Signed-off-by: Yuanjing Xue <197832395+yuanjingx87@users.noreply.github.com> Signed-off-by: Yanchao Lu <yanchaol@nvidia.com> Co-authored-by: Yanchao Lu <yanchaol@nvidia.com>	2025-10-20 09:46:44 -07:00
HuiGao-NV	d0663e16e0	[https://nvbugs/5492250 ][fix] Remove isolated cases and unwaive cases (#8492 ) Signed-off-by: Hui Gao <huig@nvidia.com>	2025-10-20 07:40:07 -04:00
Pamela Peng	b818a912d7	[https://nvbugs/5540752 ][fix] Support quantized Phi4 MM models (#8190 ) Signed-off-by: Pamela <179191831+pamelap-nvidia@users.noreply.github.com>	2025-10-20 06:36:09 -04:00
Robin Kobus	18c7a520b3	[None][feat] Update devcontainer configuration to include additional extensions (#8369 ) Signed-off-by: Robin Kobus <19427718+Funatiq@users.noreply.github.com>	2025-10-20 12:29:15 +02:00
mpikulski	97ce0ecefe	[TRTLLM-8436][feat] batched sampling and top-k logprobs improvements (#8398 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-10-20 11:15:41 +02:00
QI JUN	d05079ba4b	[None][ci] move some test cases from H100 to A10 (#8449 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>	2025-10-20 01:58:34 -04:00
Yi Zhang	3c2b3bd4d4	[TRTLLM-7255][feat] Add iteration log parser script for benchmark log (#6942 ) Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>	2025-10-20 01:34:52 -04:00
Zhanrui Sun	8124a62b74	[TRTLLM-8669][infra] Use artifactory mirror for install python (#8394 ) Signed-off-by: ZhanruiSunCh <184402041+ZhanruiSunCh@users.noreply.github.com> Signed-off-by: Zhanrui Sun <184402041+ZhanruiSunCh@users.noreply.github.com> Signed-off-by: Yanchao Lu <yanchaol@nvidia.com> Co-authored-by: Yanchao Lu <yanchaol@nvidia.com>	2025-10-20 00:17:10 -04:00
Yuxian Qiu	ec32711b1e	[https://nvbugs/5542862 ][fix] Upgrade fmha_v2. (#8364 ) Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com>	2025-10-20 10:20:23 +08:00
ChristinaZ	c8b9998acb	[TRTLLM-8637][feat] Optimize the routing kernel for DeepseekV3 (MoE CUTLASS backend); Add support for KimiK2 and Qwen-next (MoE TRTLLM backend) (#7761 ) Signed-off-by: Christina Zhang <83400082+ChristinaZ@users.noreply.github.com>	2025-10-20 10:08:31 +08:00
xiweny	f7722e2b65	[TRTLLM-4866] [test] Support waiving unit tests by waives.txt (#8359 ) Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com>	2025-10-20 09:52:51 +08:00
Yueh-Ting (eop) Chen	128a351bdc	[None][fix] Avoid overwrite of `kv_cache_config.max_tokens` for VSWA scheme for the KVCacheManager (#8219 ) For VSWA scheme, we do not want `kv_cache_cnonfig.max_token` to control and cap the maximum memory of a block pool because block pool size are not identical amongst different window sizes. This MR omits the effect of `kv_cache_config.max_tokens` under `kvCacheManager.cpp` to allow the setting of block pool size to rely on the window size to share ratio and the total gpu memory analyzed and fed to the kv cache manager. Only skipping for VSWA scheme, no extra coverage was added. Signed-off-by: eopXD <yuehtingc@nvidia.com>	2025-10-20 10:48:40 +09:00
xinhe-nv	9aa086d3bb	[None][chore] update test duration (#8377 ) Signed-off-by: Xin He (SW-GPU) <200704525+xinhe-nv@users.noreply.github.com>	2025-10-19 20:45:51 -04:00
Emma Qiao	796891ba2a	[None][infra] Skip a failed case in pre-merge for main on 10/19 (#8479 ) Signed-off-by: qqiao <qqiao@nvidia.com>	2025-10-19 22:19:00 +08:00
Bo Deng	dd25595ae8	[TRTLLM-7964][infra] Set nixl to default cache transceiver backend (#7926 ) Signed-off-by: Bo Deng <deemod@nvidia.com>	2025-10-19 19:24:43 +08:00
Emma Qiao	e185173240	[None][infra] Waive test for main branch on 10/18 (#8472 ) Signed-off-by: qqiao <qqiao@nvidia.com>	2025-10-19 04:36:42 -04:00
jthomson04	852316886e	[None][fix] Fix KV event consumption (#6346 ) Signed-off-by: jthomson04 <jwillthomson19@gmail.com>	2025-10-18 15:41:26 -07:00
brb-nv	7cc65a6296	[None][chore] Waive failing transceiver test (#8473 ) Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>	2025-10-18 17:22:10 -04:00
Lucas Liebenwein	41169fb20c	[None][feat] AutoDeploy: chunked prefill support (#8158 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-10-18 00:47:35 -07:00
QI JUN	4a8ac8dd62	[TRTLLM-8480][chore] clean create_py_executor API (#8412 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>	2025-10-17 23:52:02 -04:00
Wanli Jiang	58b43a6dab	[None][fix] Fix get_num_tokens_per_image for nano-v2-vlm (#8425 ) Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>	2025-10-18 08:51:35 +08:00
Kyle McGill	136e0e6882	[None][feat] Enable CUDA graph support for KvConnectorWorker API (#8275 ) Signed-off-by: Kyle McGill <kmcgill@nvidia.com> Signed-off-by: Kyle McGill <101670481+nv-kmcgill53@users.noreply.github.com>	2025-10-17 18:09:03 -04:00
Anish Shanbhag	5ff4f88be6	[TRTLLM-8683][chore] Migrate PluginConfig to Pydantic (#8277 ) Signed-off-by: Anish Shanbhag <ashanbhag@nvidia.com>	2025-10-17 16:13:22 -04:00
h-guo18	55fed1873c	[None][chore] AutoDeploy: cleanup old inference optimizer configs (#8039 ) Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com> Co-authored-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-10-17 15:55:57 -04:00
Grzegorz Kwasniewski	bb7fdcebf4	[TRTLLM-8201][feat] Topological graph helpers (#8457 ) Signed-off-by: greg-kwasniewski1 <213329731+greg-kwasniewski1@users.noreply.github.com>	2025-10-17 12:34:19 -04:00
Venky	8d07580c95	[None] [chore] Add ATTRIBUTIONS-{CPP,Python}.md + Update in wheels setup (#8438 ) Signed-off-by: Venky Ganesh <23023424+venkywonka@users.noreply.github.com>	2025-10-17 06:33:05 -07:00
Wanli Jiang	56f697be2e	[None][feat] Add fmha_v2 kernel for head_dim=80 and sm=100 to support VLM (#8392 ) Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>	2025-10-17 19:42:47 +08:00
xinhe-nv	bc833d3de3	[TRTLLM-8638][fix] add waives tests (#8445 ) Signed-off-by: Xin He (SW-GPU) <200704525+xinhe-nv@users.noreply.github.com>	2025-10-17 03:37:53 -07:00
Perkz Zheng	0722717ec0	[None][fix] trtllm-gen regression in PR 8301 (#8426 ) Signed-off-by: Perkz Zheng <67892460+PerkzZheng@users.noreply.github.com>	2025-10-17 03:21:31 -07:00
zhhuang-nv	7a2bab93f0	[None][test] Add post merge test for Seed-OSS-36B-Instruct (#8321 ) Signed-off-by: Zhen Huang <145532724+zhhuang-nv@users.noreply.github.com>	2025-10-17 02:30:33 -07:00
Yanchao Lu	e72ade33c2	[None][chore] Update commit msg for adding lock files (#8448 ) Signed-off-by: Yanchao Lu <yanchaol@nvidia.com>	2025-10-17 00:24:26 -07:00
Leslie Fang	023e515d33	[None][chore] Combine two documents of feature combination matrix (#8442 ) Signed-off-by: leslie-fang25 <leslief@nvidia.com>	2025-10-17 14:31:33 +08:00
yufeiwu-nv	1e1f430163	[None][test] Filter out all fp8 test case for A100. (#8420 ) Signed-off-by: yufeiwu <230315618+yufeiwu-nv@users.noreply.github.com>	2025-10-16 20:42:50 -07:00
Ivy Zhang	70a0f5beb6	[TRTLLM-8580][test] save runtime report periodically (#8312 ) Signed-off-by: Ivy Zhang <25222398+crazydemo@users.noreply.github.com>	2025-10-17 10:47:26 +08:00
Tracin	dd06612d0e	[https://nvbugs/5540138 ][fix] Fix shape error when duplicating kv. (#8390 ) Signed-off-by: Tracin <10434017+Tracin@users.noreply.github.com>	2025-10-17 10:07:29 +08:00
yuanjingx87	85deacf117	[None][infra] Update CI allowed list 2025_10_15 (#8403 ) Signed-off-by: Yuanjing Xue <197832395+yuanjingx87@users.noreply.github.com>	2025-10-16 14:17:34 -07:00
yuanjingx87	3481d03470	[None][infra] Fix for generate lockfile pipeline (#7820 ) Signed-off-by: Yuanjing Xue <197832395+yuanjingx87@users.noreply.github.com>	2025-10-16 14:17:18 -07:00
Iman Tabrizian	22eb1633ae	[None][bug] Set NCCL_GRAPH_REGISTER to false to avoid hang (#8413 ) Signed-off-by: Iman Tabrizian <10105175+tabrizian@users.noreply.github.com>	2025-10-16 18:59:18 +02:00
John Calderon	46ee7acb33	[TRTLLM-6780][fix] Add multimodal data to dummy requests during memory profiling (#7539 ) Signed-off-by: John Calderon <johncalesp@gmail.com> Signed-off-by: John Calderon <jcalderon@nvidia.com> Signed-off-by: john calderon <jcalderon@nvidia.com> Signed-off-by: John Calderon <jcalderon@nvidia>	2025-10-16 17:49:22 +02:00
Yanchao Lu	bde606f82d	Update Dockerfile.multi Signed-off-by: Yanchao Lu <yanchaol@nvidia.com>	2025-10-16 22:46:19 +08:00
Jin Li	d594c2d0ff	[https://nvbugs/5537348 ][fix] Use device tensor index for MTP (#8062 ) Signed-off-by: Jin Li <59594262+liji-nv@users.noreply.github.com> Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com>	2025-10-16 22:46:19 +08:00
Yiqing Yan	05dd437084	[https://nvbugs/5565541 ][fix] Add timeout threshold for H100 FHMA test (#8354 ) Signed-off-by: Yiqing Yan <yiqingy@nvidia.com> Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com>	2025-10-16 22:46:19 +08:00

1 2 3 4 5 ...

3260 Commits