TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-02-08 04:01:51 +08:00

Author	SHA1	Message	Date
Yukun He	39076410a8	[https://nvbugs/5676748 ][fix] Fix mismatched nvfp4 gemm sf shape. (#9336 ) Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>	2025-11-24 12:16:32 +08:00
brb-nv	c045e359a7	[https://nvbugs/5637012 ][fix] Fix helix unit tests (#9369 ) Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>	2025-11-23 19:34:22 -08:00
Yukun He	c3acf965a6	[TRTLLM-7963][fix] Several improvements of autotuning quality (#9348 ) * Skip the shape profile generating process if the profile has already been found in the cache under tuning mode. This is a prerequisite for nested autotuning because host overhead might be included during the profiling of the high-level op. * Enable the profiling with CUDA graph as the default profiling method. * Apply a heuristic method to cut off the number of repeat times of profiling according to a few-run time measurement.	2025-11-24 10:38:45 +08:00
Bo Li	fcfec93cad	[TRTLLM-9389][chore] Rename AlltoAll backend names (#9329 ) Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>	2025-11-23 13:52:57 -08:00
William Zhang	11a0b276fb	[#9230 ][feat] Slimmed down implementation of nemotron H (#9235 ) * Why? The reference nemotron H code on HuggingFace is out of date, and therefore bugged, and has several untested code paths. This makes an already hairy patching system even hairier. The proposal is to do away with those patches, and replace the original implementation with one that is heavily slimmed down. * What? This PR sets the basis for an alternative path with such a slimmed down implementation that: - fixes bugs in the current HF implementation - adds no new dependencies to TensorRT-LLM - does away with unnecessary features for TensorRT-LLM/ AutoDeploy: - no training related code (dropout, gradient checkpointing, etc.) - no caching logic (we want to replace it with our own anyway) - no attention masking where possible - reuses existing AD custom ops for mamba SSM update / causal conv1d / attention In order for the above to be usable in the AD apparatus, `AutoModelForCausalLMFactory` is extended to allow registrations of custom model implementations. Signed-off-by: William Zhang <133824995+2ez4bz@users.noreply.github.com>	2025-11-23 03:13:32 -08:00
Neta Zmora	3952a61681	[#9388 ][fix] AutoDeploy: Fix cutlass BF16 MoE kernel invocation (#9339 ) Signed-off-by: Neta Zmora <96238833+nzmora-nvidia@users.noreply.github.com>	2025-11-21 17:05:03 -08:00
Chenghao Zhang	564989865c	[TRTLLM-9082][feat] AutoDeploy: Move the moe Align kernel to AOT (#9106 ) Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>	2025-11-21 16:05:48 -08:00
Izzy Putterman	eb7792e875	[None][feat] Eagle: PostNorm and multilayer options (#9233 ) Signed-off-by: Izzy Putterman <iputterman@nvidia.com>	2025-11-21 17:39:00 -05:00
Enwei Zhu	13fbd4366a	[TRTLLM-9370][feat] Integration of CuteDSL NVFP4 grouped GEMM (Part 2: SwiGLU Fusion and Finalize Fusion) (#9288 ) Signed-off-by: Enwei Zhu <21126786+syuoni@users.noreply.github.com>	2025-11-21 14:03:38 -08:00
Ziyi Xiong	5df907b388	[https://nvbugs/5590408 ][fix] Fallback to greedy sampling in two-model overlap scheduler (#9321 ) Signed-off-by: ziyixiong-nv <219238287+ziyixiong-nv@users.noreply.github.com>	2025-11-21 10:19:59 -05:00
HuiGao-NV	6dd2fcd7b3	[https://nvbugs/5629833 ][fix] Don't fill tensors with 0 (#9296 ) Signed-off-by: Hui Gao <huig@nvidia.com>	2025-11-21 20:50:05 +08:00
mpikulski	095b6864a8	[TRTLLM-8650][fix] beam search request validation (#8433 ) (#9228 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-11-21 04:08:45 -08:00
xxi	cc0dc7c124	[TRTLLM-8957][feat] create communication related classes (#8968 )	2025-11-20 22:32:42 -08:00
Yukun He	9a79f32f7a	[https://nvbugs/5608489 ][fix] Fix output unpack issues for Llama3/4 NVFP4 models. (#8679 ) Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com> Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com> Signed-off-by: Mike Iovine <miovine@nvidia.com>	2025-11-20 12:43:13 -05:00
Lizhi Zhou	33b0b945c7	[https://nvbugs/5582277 ][fix] rework DisaggPPTerminationHandler to fix hang issue (#8519 ) Signed-off-by: Lizhi Zhou <1432185+reasonsolo@users.noreply.github.com> Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com> Signed-off-by: Mike Iovine <miovine@nvidia.com>	2025-11-20 12:43:13 -05:00
Jin Li	3454eacd74	[https://nvbugs/5546510 ][fix] Move torch.cuda.Stream out of torch com… (#8494 ) Signed-off-by: Jin Li <59594262+liji-nv@users.noreply.github.com> Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com> Signed-off-by: Mike Iovine <miovine@nvidia.com>	2025-11-20 12:43:13 -05:00
JunyiXu-nv	ee6944bfa2	[https://nvbugs/5569713 ][fix] Disable fp8 deep gemm for EXAONE-4.0-32B-FP8 (#8429 ) Signed-off-by: Junyi Xu <219237550+JunyiXu-nv@users.noreply.github.com> Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com> Signed-off-by: Mike Iovine <miovine@nvidia.com>	2025-11-20 12:43:13 -05:00
Liao Lanyu	04ad9f96fa	[https://nvbugs/5667687 ][fix] Set correct lm_head_tp_size_upper_bound (#9300 ) Signed-off-by: Lanyu Liao <lancelly@users.noreply.github.com> Co-authored-by: Lanyu Liao <lancelly@users.noreply.github.com>	2025-11-20 00:41:00 -08:00
Neta Zmora	1d6fbbf45d	[#9236 ][feature] Make sharing of activation_type across SW layers more robust (#9238 ) C++, Python and Python MoE layer all share the definition of ActivationType. Currently this is done thru redefinition which is fragile and can break when adding new activation function types. tensorrt_llm/_torch/utils.py cpp/tensorrt_llm/kernels/cutlass_kernels/include/common.h => tensorrt_llm/layers/moe.py cpp/tensorrt_llm/kernels/cutlass_kernels/moe_gemm/moe_kernels.cu Signed-off-by: Kaiyu Xie <26294424+kaiyux@users.noreply.github.com> Signed-off-by: Neta Zmora <96238833+nzmora-nvidia@users.noreply.github.com> Co-authored-by: Kaiyu Xie <26294424+kaiyux@users.noreply.github.com>	2025-11-20 16:06:58 +08:00
Yechan Kim	d5622b2689	[None][fix] Multimodal InputProcessor dummy builder fix (#8916 ) Signed-off-by: yechank <161688079+yechank-nvidia@users.noreply.github.com>	2025-11-19 22:32:21 -08:00
Chang Liu	79a6c9742b	[None][fix] Use fp32 for indexer weight_proj GEMM (#9243 ) Signed-off-by: Chang Liu (Enterprise Products) <9713593+chang-l@users.noreply.github.com>	2025-11-19 21:52:38 -08:00
Neta Zmora	028fc877a5	[#9096 ][feature] Auto Deploy: configurable fused MoE backend (#9194 ) Allow configuring Auto Deploy's MoE/FP8-MoE backend from external yaml config file. Signed-off-by: Neta Zmora <96238833+nzmora-nvidia@users.noreply.github.com>	2025-11-19 21:50:22 -08:00
Yukun He	b6bced83c0	[TRTLLM-7963][feat] Use CUDAGraph to improve the tuning accuracy for AutoTuner. (#9089 ) Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>	2025-11-20 08:54:29 +08:00
Fanrong Li	d4abb86f3e	[None][fix] fix EPLB for DeepSeek-V3.2-Exp (#9245 ) Signed-off-by: Fanrong Li <23290157+lfr-0531@users.noreply.github.com>	2025-11-19 13:45:54 -08:00
Faraz	49c45ebef1	[None][fix] change logging for weight loading on unified memory (#9177 ) Signed-off-by: Faraz Khoubsirat <58580514+farazkh80@users.noreply.github.com> Signed-off-by: Simeng Liu <109828133+SimengLiu-nv@users.noreply.github.com> Co-authored-by: Simeng Liu <109828133+SimengLiu-nv@users.noreply.github.com>	2025-11-19 14:31:19 -05:00
NVShreyas	1eae941d77	[#9237 ][feat] enable iter stats in autodeploy (#9278 ) Signed-off-by: Shreyas Misra <shreyasm@nvidia.com>	2025-11-19 19:29:29 +01:00
NVShreyas	a7c0b54ce7	[None][feat] add specdec to nemotron nas (#8985 ) Signed-off-by: Shreyas Misra <shreyasm@nvidia.com>	2025-11-19 19:28:35 +01:00
Bo Li	d8b05894ee	[None][perf] Adjust select_alltoall_method_type. (#8950 ) Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>	2025-11-19 07:43:55 -08:00
mpikulski	46dd9886bb	[https://nvbugs/5661877 ][fix] fix test regression in TestBatchedSampling::test_samples (#9215 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-11-19 01:44:44 -08:00
CarstyYou	ee941ac779	[https://nvbugs/5456493 ][feat] add fp8 dense for sm120 (#9174 ) Signed-off-by: CarstyYou <186021327+CarstyYou@users.noreply.github.com>	2025-11-19 14:40:34 +08:00
ChristinaZ	941a54c66a	[None][feat] Update the indexer topK (#9255 ) Signed-off-by: Christina Zhang <83400082+ChristinaZ@users.noreply.github.com>	2025-11-19 11:49:00 +08:00
jellysnack	99ba723e20	[None][fix] logits device and shape issues in dynamic draft path (#9079 ) Signed-off-by: jellysnack <oleg.jellysnack@gmail.com>	2025-11-18 19:22:47 -08:00
Grzegorz Kwasniewski	7905d6c0da	[#9098 ][feat] Simple sharding latent experts (#9099 ) Signed-off-by: greg-kwasniewski1 <213329731+greg-kwasniewski1@users.noreply.github.com>	2025-11-18 21:14:22 -05:00
Grzegorz Kwasniewski	92f86a50d4	[#9137 ][feat] Factory sharding as default (#9144 ) Signed-off-by: greg-kwasniewski1 <213329731+greg-kwasniewski1@users.noreply.github.com>	2025-11-18 21:12:03 -05:00
Patrice Castonguay	9b0f45298f	[None][feat] Have ability to cancel disagg request if KV cache resource are exhausted (#9155 ) Signed-off-by: Patrice Castonguay <55748270+pcastonguay@users.noreply.github.com>	2025-11-18 20:59:17 -05:00
Enwei Zhu	7c4777a571	[TRTLLM-9286][feat] Integration of CuteDSL NVFP4 grouped GEMM (#8880 ) Signed-off-by: Enwei Zhu <21126786+syuoni@users.noreply.github.com>	2025-11-18 17:40:12 -08:00
Ziyi Xiong	7c4344b92e	[https://nvbugs/5590408 ][fix] Exclude num of draft tokens from mMaxSeqLenKv (#9210 ) Signed-off-by: ziyixiong-nv <219238287+ziyixiong-nv@users.noreply.github.com>	2025-11-18 15:41:56 -05:00
Eran Geva	3ac11a6180	[#9152 ][fix] AutoDeploy fused_allreduce_residual_rmsnorm to support demollm mode (#9197 ) Signed-off-by: Eran Geva <19514940+MrGeva@users.noreply.github.com>	2025-11-18 22:15:29 +02:00
Chenghao Zhang	f0b68e4c66	[None][feat] AutoDeploy: Perf improvement for small batch size (#9163 ) Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com> Co-authored-by: Suyog Gupta <41447211+suyoggupta@users.noreply.github.com>	2025-11-18 12:11:12 -08:00
Zheyu Fu	c4e02d7f04	[TRTLLM-8136][feat] Dynamic draft length in spec decode (stage 1). (#8194 ) Signed-off-by: Zheyu Fu <zheyuf@NVIDIA.com>	2025-11-18 11:13:39 -05:00
Robin Kobus	9913dc25ae	[None][refactor] decoding inputs, part 2 (#5799 ) Signed-off-by: Robin Kobus <19427718+Funatiq@users.noreply.github.com>	2025-11-18 14:38:51 +01:00
Chang Liu	8e001dd195	[None][fix] DeepSeek V3.2 indexer RoPE fix (#9232 ) Signed-off-by: Chang Liu (Enterprise Products) <9713593+chang-l@users.noreply.github.com>	2025-11-18 20:35:27 +08:00
Lizhi Zhou	07343bb11c	[None][chore] fix a deepseekv3 error when debug mode is on (#9217 ) Signed-off-by: Lizhi Zhou <1432185+reasonsolo@users.noreply.github.com>	2025-11-18 01:14:32 -08:00
ruodil	82480346aa	[https://nvbugs/5652552 ][fix] add printing for llm args (#9205 ) Signed-off-by: Ruodi Lu <ruodil@users.noreply.github.com> Co-authored-by: Ruodi Lu <ruodil@users.noreply.github.com>	2025-11-17 23:58:36 -08:00
Tri Dao	fc088e642c	[None][feat] Support Glm4MoeForCausalLM (#8256 ) Signed-off-by: Tri Dao <daominhtri0503@gmail.com> Co-authored-by: Xuanyu Chen <xuanyuc@nvidia.com>	2025-11-18 09:43:21 +08:00
Robin Kobus	df41f220a2	[TRTLLM-8831][feat] Enable early exit with overlap scheduler (#8587 ) Signed-off-by: Robin Kobus <19427718+Funatiq@users.noreply.github.com>	2025-11-17 18:07:13 +01:00
Mike Iovine	6151a4c9d6	[None][feat] Add simple optimizations for MTP 2-model (#9176 ) Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com>	2025-11-17 10:05:39 -05:00
Kaiyu Xie	04be5a704e	[None] [fix] Fix missing ActivationType issue (#9171 ) Signed-off-by: Kaiyu Xie <26294424+kaiyux@users.noreply.github.com> Signed-off-by: Enwei Zhu <21126786+syuoni@users.noreply.github.com> Signed-off-by: Neta Zmora <96238833+nzmora-nvidia@users.noreply.github.com> Co-authored-by: Enwei Zhu <21126786+syuoni@users.noreply.github.com> Co-authored-by: Neta Zmora <96238833+nzmora-nvidia@users.noreply.github.com>	2025-11-17 10:43:25 +08:00
Anthony Chang	86cfb3ea7e	[None][feat] Update TRTLLM MoE cubins; reduce mxfp4 weight padding requirement; tighten TMA bound (#9025 ) Signed-off-by: Anthony Chang <27950904+rosenrodt@users.noreply.github.com>	2025-11-17 10:04:29 +08:00
Jinyang Yuan	6dc70aa0e5	[https://nvbugs/5613089 ][fix] Fix the rank to access all_rank_chunk_size_list when chunked MoE is used (#8723 ) Signed-off-by: Jinyang Yuan <154768711+jinyangyuan-nvidia@users.noreply.github.com>	2025-11-17 10:01:08 +08:00
sunnyqgg	7862b15a65	[TRTLLM-8778][feat] Add tree attention support for blackwell arch (#8975 ) Signed-off-by: qgai <qgai@nvidia.com>	2025-11-17 09:01:53 +08:00
Guoming Zhang	e0f69657c7	[None][fix] Update the attention layers counting for Qwen3-next. (#9072 ) Signed-off-by: nv-guomingz <137257613+nv-guomingz@users.noreply.github.com>	2025-11-16 11:52:56 -08:00
JadoTu	3cde84581d	[None][fix] Make the sliced nvfp4 output contiguous (#9123 ) Signed-off-by: jiant <107457950+JadoTu@users.noreply.github.com>	2025-11-15 20:00:54 +08:00
Chenghao Zhang	f6f6e1f25d	[#9102 ][feat] AutoDeploy: Support fp8 kv cache (#9107 ) Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>	2025-11-13 23:55:45 -08:00
Lizhi Zhou	8bd779171e	[https://nvbugs/5631254 ][fix] avoid torch.compile for multiple times (#9135 ) Signed-off-by: Lizhi Zhou <1432185+reasonsolo@users.noreply.github.com>	2025-11-13 21:49:52 -08:00
Suyog Gupta	d12cb9436d	[None][feat] Autodeploy add triton configs and optimize mamba prefill (#9083 ) Signed-off-by: Suyog Gupta <41447211+suyoggupta@users.noreply.github.com>	2025-11-13 19:15:43 -08:00
heyuhhh	f07e9977c6	[None] [feat] Use triton kernels for RocketKV prediction module (#8682 ) Signed-off-by: yuhangh <58161490+heyuhhh@users.noreply.github.com>	2025-11-13 18:51:09 -08:00
Neta Zmora	34dc6869f3	[#8732 ][feat] Update TRTLLM Cutlass MoE kernels with ReLU2 (#9011 ) Update TRTLLM Cutlass MoE kernels with ReLU2 activation. Nemotron-6 requires ReLU2 (i.e. squared ReLU) MoE activation function. The PR adds this and adds an API to set the activation function, in general. The ReLU2 changes are based on this FlashInfer PR: https://github.com/flashinfer-ai/flashinfer/pull/1954. The PR also updates the Auto Deploy MoE backend for 16-bit and FP8 from Triton (`torch.ops.auto_deploy.triton_moe_fused`, `torch.ops.auto_deploy.triton_quant_fp8_moe`) to TRTLLM/Cutlass (`torch.ops.auto_deploy.trtllm_moe_fused`, `torch.ops.auto_deploy.trtllm_quant_fp8_moe_fused`). Signed-off-by: Neta Zmora <96238833+nzmora-nvidia@users.noreply.github.com> Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com> Co-authored-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>	2025-11-13 16:54:45 -08:00
dongxuy04	a370643b26	[None][fix] support topk autotuner input for expert slot per group larger than 32 (#9087 ) Signed-off-by: Dongxu Yang <78518666+dongxuy04@users.noreply.github.com>	2025-11-14 08:37:20 +08:00
Leslie Fang	daa31d78f4	[https://nvbugs/5652552 ][fix] Log the llm args for main branch (#9120 ) Signed-off-by: leslie-fang25 <leslief@nvidia.com>	2025-11-14 07:43:21 +08:00
Frida Hou	b51258acdd	[None][autodeploy] fix weight extraction for graph based quantized checkpoints (#9109 ) Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>	2025-11-13 13:14:24 -08:00
Frida Hou	e96a3d294d	[None][autodeploy] minor refactor to rmsnorm transforms (#8657 ) Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>	2025-11-13 13:13:58 -08:00
Jinyang Yuan	12f339f3bf	[None][fix] Fix the aux_stream in Llama4MinLatencyFusedMoE (#9035 ) Signed-off-by: Jinyang Yuan <154768711+jinyangyuan-nvidia@users.noreply.github.com>	2025-11-13 09:09:52 -08:00
Ziyi Xiong	a7aaf50541	[TRTLLM-8084][feat] Enhance the overlap shceduler for two-model spec decoding (#8706 ) Signed-off-by: ziyixiong-nv <219238287+ziyixiong-nv@users.noreply.github.com>	2025-11-13 10:20:16 -05:00
Chang Liu	c37924f37b	[None][fix] Clear indexer k cache reference before release cuda memory (#9110 ) Signed-off-by: Chang Liu (Enterprise Products) <9713593+chang-l@users.noreply.github.com>	2025-11-12 22:12:53 -08:00
Zhang Ge	49df731b96	[#6507 ][fix] Fix precision issue due to KV layout mismatch for split/concat kernels (#6917 ) Signed-off-by: ZhangGe6 <sjtu.zg123@gmail.com> Co-authored-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com>	2025-11-13 12:14:58 +08:00
QI JUN	d1b003d31e	[TRTLLM-9212][chore] move MoeLoadBalancerConfig to llm_args.py (#9002 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>	2025-11-13 10:47:35 +08:00
Chenghao Zhang	f1d637ec69	[None][fix] AutoDeploy: Use tmp folder for the load_moe_align (#9101 ) Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>	2025-11-12 14:59:49 -08:00
dongxuy04	9241ccaf27	[None][feat] Enable EPLB for trtllm-gen and cutlass backend (#8886 ) Signed-off-by: Dongxu Yang <78518666+dongxuy04@users.noreply.github.com>	2025-11-12 12:30:27 -08:00
Patrice Castonguay	8a751a0e56	[None][chore] Remove is_disaggregated param in executor request queue (#9049 ) Signed-off-by: Patrice Castonguay <55748270+pcastonguay@users.noreply.github.com>	2025-11-12 13:37:15 -05:00
Fanrong Li	780d4f9dc5	[None][feat] Add MTP>1 support for DS-v3.2 (#9045 ) Signed-off-by: Fanrong Li <23290157+lfr-0531@users.noreply.github.com>	2025-11-12 09:56:12 -08:00
Neta Zmora	53491ffdb1	[#9023 ][feat] reduce AD graph optimization time for non-participating passes (#9024 ) Shorten AD graph optimization by 30% (measured on Nemotron-6): A bug in the transformation interface marked all passes as not clean, regardless of what was reported by the transformation Fix how the optimization passes report the results of their actions. Many passes report that the graph is not clean even when they didn't participate in the optimization. Each graph cleaning invocation can take several seconds. Signed-off-by: Neta Zmora <96238833+nzmora-nvidia@users.noreply.github.com>	2025-11-12 09:05:53 -08:00
Chang Liu	0b81173efa	[TRTLLM-9259][perf] Use torch.compile to fuse copy + layernorm within the LayerNorm module (#9052 ) Signed-off-by: Chang Liu <9713593+chang-l@users.noreply.github.com>	2025-11-11 18:11:00 -08:00
Lucas Liebenwein	aca56097cb	[None][fix] AutoDeploy: update nano3 accuracy test (#9061 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-11-11 12:26:31 -08:00
QI JUN	524754b6fd	[TRTLLM-8521][chore] remove circular dependency between model engine and cuda graph runner (#7572 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>	2025-11-11 10:13:45 -08:00
Chenghao Zhang	ec9cf715a2	[None][feat] AutoDeploy: Perf improvement for mamba layers (#8991 ) Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com> Signed-off-by: Suyog Gupta <41447211+suyoggupta@users.noreply.github.com> Co-authored-by: Suyog Gupta <41447211+suyoggupta@users.noreply.github.com>	2025-11-11 08:27:07 -08:00
Wanli Jiang	ebdd1cc8e0	[TRTLLM-8119][feat] Update doc/tests/chat_template for nano-v2-vlm (#8840 ) Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>	2025-11-11 07:48:23 -08:00
mpikulski	20fd305bb6	[None][fix] type annotation (#9071 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-11-11 07:20:20 -08:00
mpikulski	b151de4a8f	[TRTLLM-8377][test] unit tests for TorchSampler batched sampling (#9012 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com> Co-authored-by: Robin Kobus <19427718+Funatiq@users.noreply.github.com>	2025-11-11 07:16:42 -08:00
Guoming Zhang	b894dc2d70	[None][fix] Display the GPU memory information in GiB unit. (#9070 ) Signed-off-by: nv-guomingz <137257613+nv-guomingz@users.noreply.github.com>	2025-11-11 06:24:59 -08:00
mpikulski	979b3ae9ce	[TRTLLM-7723][feat] sampling using FlashInfer.sampling (#8581 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-11-11 03:21:19 -08:00
Yuxian Qiu	7aeac97e4e	[https://nvbugs/5622938 ][fix] Use async send_requests_to_next_pp. (#9041 ) Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com>	2025-11-11 14:19:44 +08:00
Lucas Liebenwein	6bf4e59267	[#8763 ][feature] AutoDeploy: configurable dtype for caching (#8812 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-11-10 22:17:14 -08:00
Chang Liu	7ceb5e5ab6	[TRTLLM-9198][perf] Add torch.compile + multi-stream support for k-cache scatter and weight scaling (#8988 ) Signed-off-by: Chang Liu (Enterprise Products) <9713593+chang-l@users.noreply.github.com> Co-authored-by: Fanrong Li <23290157+lfr-0531@users.noreply.github.com>	2025-11-11 12:33:30 +08:00
Frida Hou	f40e1f7496	[https://nvbugs/5625972 ][fix] Add context manager to fix FakeTensorProp (#9047 ) Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>	2025-11-10 16:25:58 -08:00
mpikulski	edc91ba819	[None][fix] Improve type annotations on ResourceManager.get_resource_manager (#9013 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-11-10 15:06:16 +01:00
ChristinaZ	2e7769d1e8	[None][feat] Add customized topk and related unit tests for DSA (#8882 ) Signed-off-by: Christina Zhang <83400082+ChristinaZ@users.noreply.github.com>	2025-11-10 03:35:35 -08:00
bhsueh_NV	e8d4a56dd0	[None][fix] fix eagle3 accuracy issue on sm120 (#8944 ) Signed-off-by: bhsueh <11360707+byshiue@users.noreply.github.com>	2025-11-10 14:02:03 +08:00
Fanrong Li	a7033a9193	[TRTLLM-9001][feat] add TP support for DeepSeek-V3.2 (#8943 ) Signed-off-by: Fanrong Li <23290157+lfr-0531@users.noreply.github.com>	2025-11-10 12:16:01 +08:00
Chang Liu	7081f254cf	[None][perf] Add custom indexer k cache scatter op (#8960 ) Signed-off-by: Chang Liu (Enterprise Products) <9713593+chang-l@users.noreply.github.com>	2025-11-07 11:24:26 -08:00
Patrice Castonguay	d8ea0b967f	[None][fix] Moving transfer timeout test to test_llm_pytorch, fixing broken kv transfer timeout (#8892 ) Signed-off-by: Patrice Castonguay <55748270+pcastonguay@users.noreply.github.com>	2025-11-07 07:33:51 -08:00
mpikulski	5ef65872a3	[None][fix] type annotations in fuse_input_embeds (#8976 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-11-07 09:04:08 +01:00
Stefan Niebler	326a201473	[https://nvbugs/5508536 ][fix] Take Over (#8627 ): Reintroduce: Move stop_criteria to sample_async (#7041 ) (#8794 ) Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com> Signed-off-by: Stefan Niebler <82932102+stnie@users.noreply.github.com> Co-authored-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>	2025-11-07 09:01:15 +01:00
QI JUN	1c6e490894	[TRTLLM-9065][chore] remove PyTorchConfig completely (#8856 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>	2025-11-06 22:37:03 -08:00
Eran Geva	990e674b71	[None][fix] Switch AD AllReduce strategy to NCCL (#8979 ) Signed-off-by: Eran Geva <19514940+MrGeva@users.noreply.github.com>	2025-11-07 06:49:44 +02:00
xiweny	ee20e679a9	[https://nvbugs/5636986 ][fix] Fix DeepGemmMoe get_buffer calls (#8939 ) Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Signed-off-by: xiweny <13230610+VALLIS-NERIA@users.noreply.github.com>	2025-11-06 19:57:19 -08:00
Cao Dong	b53961e972	[None][feat] Return logprobs incrementally in torch backend (#8785 ) Signed-off-by: Dong Cao <docao@nvidia.com>	2025-11-07 10:23:39 +08:00
Chang Liu	1c19fd6868	[https://nvbugspro.nvidia.com/bug/5637012 ][fix] Bugfix when config is None for MLA (#8978 ) Signed-off-by: Chang Liu (Enterprise Products) <9713593+chang-l@users.noreply.github.com>	2025-11-07 09:37:19 +08:00
jthomson04	fcae852cef	[None][fix] Fix KV cache clearing with KV Connector API (#8750 ) Signed-off-by: jthomson04 <jwillthomson19@gmail.com>	2025-11-06 14:28:27 -08:00
Chenghao Zhang	1a78e7a3d6	[None][feat] AutoDeploy: Support Latent MOE for Nemotron (#8955 ) Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>	2025-11-06 12:40:19 -08:00
Chenghao Zhang	ddf2d010e2	[TRTLLM-8814][feat] AutoDeploy: Use TRTLLM kernels for FP8 linear (#8820 ) Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com> Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com> Signed-off-by: nvchenghaoz <211069071+nvchenghaoz@users.noreply.github.com> Co-authored-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-11-06 11:00:10 -08:00
yunruis	51545560da	[TRTLLM-8803][feat] Add rope and uk-bgemm overlap for mla generation (#8495 ) Signed-off-by: yunruis <205571022+yunruis@users.noreply.github.com>	2025-11-06 17:39:57 +08:00
JadoTu	6bbb43f2b9	[None][feat] Add qwen3-next nvfp4 support (#8526 ) Signed-off-by: jiant <107457950+JadoTu@users.noreply.github.com>	2025-11-06 09:45:44 +08:00
Frida Hou	fb7f9831d3	[#8924 ][fix] Fix AutoDeploy pattern matcher for torch 2.9 (#8920 ) Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>	2025-11-05 13:29:20 -08:00
Lucas Liebenwein	b181568d6f	[TRTLLM-8201][feat] Nemotron H MoE Sharding (#8744 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com> Signed-off-by: greg-kwasniewski1 <213329731+greg-kwasniewski1@users.noreply.github.com> Co-authored-by: greg-kwasniewski1 <213329731+greg-kwasniewski1@users.noreply.github.com> Co-authored-by: Suyog Gupta <41447211+suyoggupta@users.noreply.github.com>	2025-11-05 12:35:29 -08:00
Chang Liu	e57d83c5dc	[TRTLLM-8768][chore] Fuse QK down_proj with indexer K + weight_proj for FP4 ckpt (#8771 )	2025-11-05 07:57:09 -08:00
Yukun He	b9e5315dfb	[https://nvbugs/5623960 ][fix] Fix the logger once key issue and further compress log in AutoTuner. (#8873 ) Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>	2025-11-05 15:25:43 +08:00
Shiyu Li	eeb56c2848	[None][feat] MNNVLAllreduce Kernel Refactor (#8018 ) Signed-off-by: Shiyu Li <timlee0212@outlook.com> Co-authored-by: Jin Li <59594262+liji-nv@users.noreply.github.com>	2025-11-05 08:49:47 +08:00
Frida Hou	11ded113cd	[#8389 ][fix] Update group attention matching to first map to custom torch attention (#8638 ) Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>	2025-11-04 12:00:43 -08:00
shuyixiong	70e4d72ffa	[TRTLLM-8511][feat] Add update_weights and sleep_wakeup support for rl integration (#8302 ) Signed-off-by: shuyix <219646547+shuyixiong@users.noreply.github.com> Co-authored-by: Liwei Ma <liweim@nvidia.com> Co-authored-by: Jonas Yang CN <joyang@nvidia.com>	2025-11-04 10:19:24 -08:00
Bo Li	e4bf29bc66	[None][feat] Integrate MnnvlThroughput into TRTLLM MoE. (#8728 ) Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>	2025-11-04 21:36:29 +08:00
Cao Dong	dddfcdd3bf	[None][fix] Fix bug of undefined py_topk_logprobs_vals (#8789 ) Signed-off-by: Dong Cao <docao@nvidia.com>	2025-11-04 19:32:59 +08:00
CarstyYou	4296c9553d	[TRTLLM-1234][feat] Add fp8 blockscaled Gemm for sm120 (#8844 ) Signed-off-by: CarstyYou <186021327+CarstyYou@users.noreply.github.com>	2025-11-04 18:10:36 +08:00
danielafrimi	2b58dba0f6	[https://nvbugs/5524714 ][fix] Fix TP sharding of fused-QKV weight scales in W4A16 AWQ (#8432 ) Signed-off-by: Daniel Afrimi <dafrimi@nvidia.com> Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com>	2025-11-04 16:42:31 +08:00
Patrice Castonguay	65c138108e	[https://nvbugs/5552889 ][fix] fix: Prevent empty batch when using attention DP with disagg (#8372 ) Signed-off-by: Patrice Castonguay <55748270+pcastonguay@users.noreply.github.com> Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com>	2025-11-04 16:42:31 +08:00
xiweny	fcac2022e2	[https://nvbugs/5565565 ] [fix] fp8 wideep support sm103 (#8228 ) Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com>	2025-11-04 16:42:31 +08:00
Yechan Kim	67208f1512	[None][fix] InputProcessor config naming convention fix (#8705 ) Signed-off-by: yechank <161688079+yechank-nvidia@users.noreply.github.com>	2025-11-03 22:29:21 -08:00
HuiGao-NV	97674c3114	[TRTLLM-8690][feat] add more tensors to share buffers (#8691 ) Signed-off-by: Hui Gao <huig@nvidia.com>	2025-11-03 21:08:01 -08:00
Yan Chunwei	ed297d7c2e	[None][chore] Optimize perf for the RPC executor and add some profile utilities to llm-api (#8415 ) Signed-off-by: Superjomn <328693+Superjomn@users.noreply.github.com>	2025-11-03 17:59:49 -08:00
Matthias Jouanneaux	d0f107e4dd	[TRTLLM-5966][feat] Helix: add full MLA support for Helix (#8104 ) Signed-off-by: Matthias Jouanneaux <mjoux@nvidia.com>	2025-11-04 09:06:58 +08:00
Li Min	89336fbf07	[None][fix] Fix cute dsl nvfp4 gemm autotune issue (#8761 ) Signed-off-by: Mindy Li <11663212+limin2021@users.noreply.github.com> Co-authored-by: Kaiyu Xie <26294424+kaiyux@users.noreply.github.com>	2025-11-03 22:55:45 +08:00
Yechan Kim	f48968b6cc	[TRTLLM-6928][fix] Refactor multimodal unittest (#8453 ) Signed-off-by: yechank <161688079+yechank-nvidia@users.noreply.github.com>	2025-11-03 06:01:07 -08:00
Eran Geva	f8778230e3	[#8781 ][fix] Cache the AllReduce wrapper to avoid re-allocating workspace which caused a hang (#8803 ) Signed-off-by: Eran Geva <19514940+MrGeva@users.noreply.github.com>	2025-11-02 15:30:39 +02:00
Bo Li	4c5a8f4ec6	[None][fix] Rename: slot_count -> invalid_expert_id (#8783 ) Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>	2025-11-01 21:36:59 +08:00
QI JUN	89e0117097	[TRTLLM-8836][chore] Create ModelEngine from LlmArgs (#8600 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>	2025-11-01 05:26:06 -07:00
Fanrong Li	f0dc746738	[TRTLLM-8541][feat] Add trtllm-gen sparse MLA kernels to support per-Tensor FP8 KV Cache (#8692 ) Signed-off-by: Perkz Zheng <67892460+PerkzZheng@users.noreply.github.com> Signed-off-by: Tracin <10434017+Tracin@users.noreply.github.com> Signed-off-by: Fanrong Li <23290157+lfr-0531@users.noreply.github.com> Co-authored-by: Perkz Zheng <67892460+PerkzZheng@users.noreply.github.com> Co-authored-by: Tracin <10434017+Tracin@users.noreply.github.com>	2025-10-31 14:38:31 -07:00
Suyog Gupta	3d0e38e074	[None][perf] AutoDeploy optimize _get_unique_value (#8822 ) Signed-off-by: Suyog Gupta <41447211+suyoggupta@users.noreply.github.com>	2025-10-31 04:57:10 -07:00
Anthony Chang	852e5060aa	[https://nvbugs/5558117 ][fix] Allow per-layer quant config from hf_quant_config.json (#8617 ) Signed-off-by: Anthony Chang <27950904+rosenrodt@users.noreply.github.com>	2025-10-31 04:41:44 -07:00
Yukun He	1d4a186ace	[https://nvbugs/5623960 ][fix] Compress the warning log of AutoTuner when encountering tactic failures. (#8793 ) Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>	2025-10-31 11:09:14 +08:00
Yuxian Qiu	025d2926df	[https://nvbugs/5599515 ][fix] Fix PP bubbles. (#8687 ) Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com>	2025-10-31 10:13:56 +08:00
Chenghao Zhang	71c5576a44	[TRTLLM-8734][feat] AutoDeploy: Enable the nvfp4 for Nemotron MOE (#8737 ) Signed-off-by: nvchenghaoz <211069071+nvchenghaoz@users.noreply.github.com> Co-authored-by: Suyog Gupta <41447211+suyoggupta@users.noreply.github.com>	2025-10-30 12:33:08 -07:00
Anthony Chang	f666ad2f6b	[None][feat] Autotuner can iterate through all tactics for test purposes (#8663 ) Signed-off-by: Anthony Chang <27950904+rosenrodt@users.noreply.github.com>	2025-10-30 13:11:25 +01:00
Void	6b755fd9f8	[None][fix] fix runtime error that bf16 input is not quantized to nvfp4 when use bf16 dispatch (#8507 ) Signed-off-by: Yilin Zhang <18275976+yilin-void@users.noreply.github.com>	2025-10-30 15:06:54 +08:00
Simeng Liu	834a780655	[https://nvbugs/5599086 ][fix] Fix FP8 Linear module for spark (#8707 ) Signed-off-by: Simeng Liu <simengl@nvidia.com>	2025-10-29 13:58:19 -07:00
Iman Tabrizian	ae6875fe10	[TRTLLM-8976][feat] Move indexer-k-cache to KVCacheManager (#8699 ) Signed-off-by: Iman Tabrizian <10105175+tabrizian@users.noreply.github.com>	2025-10-29 08:04:26 -07:00
Leslie Fang	451959c60d	[TRTLLM-8763][chore] Deprecate pybind based GuidedDecodingConfig usage in torch backend (#8717 ) Signed-off-by: leslie-fang25 <leslief@nvidia.com>	2025-10-29 20:37:14 +08:00
Fanrong Li	a21697ead9	[None][fix] fix config loading for DeepSeek-V3.2 in trtllm-bench (#8729 ) Signed-off-by: Fanrong Li <23290157+lfr-0531@users.noreply.github.com>	2025-10-29 05:17:16 -07:00
kris1025	e2c5a38879	[https://nvbugs/5534574 ][fix] disable spec decoding forever once the request spec decoding is disabled (#8446 ) Signed-off-by: linquanh <linquanh@nvidia.com>	2025-10-29 19:28:43 +08:00
Yi Zhang	a69bd2a6fa	[https://nvbugs/5550409 ][fix] Disable torch compile in piecewise attention part to Avoid host overhead (#8708 ) Signed-off-by: yizhang-nv <187001205+yizhang-nv@users.noreply.github.com>	2025-10-29 18:12:58 +08:00
Chang Liu	5f737b8dbe	[None][perf] Use fp8 quant kernel in DS3.2 indexer module (#8701 ) Signed-off-by: Chang Liu (Enterprise Products) <9713593+chang-l@users.noreply.github.com>	2025-10-29 12:45:09 +08:00
Cheng Hang	15c293a90b	[None][feat] Enable nvfp4 cuda core for sm120 (#8620 ) Signed-off-by: Cheng Hang <chang@nvidia.com>	2025-10-29 12:39:03 +08:00
Yechan Kim	bc26f4ce7c	[https://nvbugs/5549829 ][fix] Qwen2.5-VL TP > 1 + Quantized weight load fix (#8680 ) Signed-off-by: yechank <161688079+yechank-nvidia@users.noreply.github.com>	2025-10-29 13:38:42 +09:00
Yechan Kim	cf8a1d2ef9	[https://nvbugs/5596377 ][fix] Fix mm dummy calculation (#8498 ) Signed-off-by: yechank <161688079+yechank-nvidia@users.noreply.github.com>	2025-10-29 09:45:21 +09:00
Kaiyu Xie	227c288441	[TRTLLM-8827] [feat] Enable low precision alltoall for Cutlass and TRTLLMGen backends (#8675 ) Signed-off-by: Kaiyu Xie <26294424+kaiyux@users.noreply.github.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>	2025-10-29 07:56:48 +08:00
Mike Iovine	00161b315f	[https://nvbugs/5549111 ][fix] Fix 2-model overlap scheduler accuracy on very long prompts (#8076 ) Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com> Signed-off-by: Michael Iovine <miovine@nvidia.com>	2025-10-28 14:55:34 -07:00
Lucas Liebenwein	0ee71d95ec	[https://nvbugs/5606166 ][fix] AutoDeploy: use tuples for cudagraph shape lookup (#8658 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-10-28 10:52:43 -07:00
Anish Shanbhag	a09b38a862	[TRTLLM-8684][chore] Migrate BuildConfig to Pydantic, add a Python wrapper for KVCacheType enum (#8330 ) Signed-off-by: Anish Shanbhag <ashanbhag@nvidia.com>	2025-10-28 09:17:26 -07:00
William Zhang	cdc9e5e645	[None][fix] Properly raise error for nemotron H models (#8697 ) Signed-off-by: William Zhang <133824995+2ez4bz@users.noreply.github.com>	2025-10-28 08:59:42 -07:00
Eran Geva	e051a05e6c	[#8694 ][fix] fix AutoDeploy cuda memory access failure in nvidia/NVIDIA-Nemotron-Nano-31B-A3-v3 (#8696 ) Signed-off-by: Eran Geva <19514940+MrGeva@users.noreply.github.com>	2025-10-28 13:21:43 +02:00
gramnarayan	88b0fbc8ff	[#8245 ][feat] Autodeploy: Guided Decoding Support (#8551 ) Signed-off-by: William Zhang <133824995+2ez4bz@users.noreply.github.com> Signed-off-by: Govind Ramnarayan <105831528+govind-ramnarayan@users.noreply.github.com> Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com> Co-authored-by: William Zhang <133824995+2ez4bz@users.noreply.github.com> Co-authored-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-10-28 09:29:57 +08:00
Bo Li	9c4432f8a4	[TRTLLM-7318][feat] MnnvlThroughput AlltoAll implementation. (#7499 ) Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com> Co-authored-by: Jin Li <59594262+liji-nv@users.noreply.github.com>	2025-10-27 13:23:06 -04:00
Chenghao Zhang	b9b2802599	[None][feat] Autodeploy: Update the ssm to use slice (#8667 ) Signed-off-by: nvchenghaoz <211069071+nvchenghaoz@users.noreply.github.com>	2025-10-27 09:45:20 -07:00
mpikulski	7c8ba71b49	[TRTLLM-8832][feat] fully async _select_generated_logits with tests (#8628 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-10-27 16:15:32 +01:00
QI JUN	4fd58137a1	[TRTLLM-8933][chore] remove unused update_executor_config function (#8678 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>	2025-10-27 10:00:47 -04:00
Kaiyu Xie	c9b08790c2	[None] [test] Add MNNVL AlltoAll tests to pre-merge (#8601 )	2025-10-27 21:39:44 +08:00
Tailing Yuan	858d6437c1	[None][fix] Fix ModelConfig.from_pretrained get quant config file (#8647 ) Signed-off-by: Tailing Yuan <yuantailing@gmail.com>	2025-10-27 11:02:24 +08:00
Jinyang Yuan	0a0f93d4a8	[None][fix] Fix the performance issue of FP8 blockwise grouped GEMM when using attention DP (#8501 ) Signed-off-by: Jinyang Yuan <154768711+jinyangyuan-nvidia@users.noreply.github.com>	2025-10-27 10:18:19 +08:00
Chenghao Zhang	a6d20f6f9b	[None][feat] AutoDeploy: Add FP8 MOE for Nemotron (#8599 ) Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> Signed-off-by: nvchenghaoz <211069071+nvchenghaoz@users.noreply.github.com> Co-authored-by: Suyog Gupta <41447211+suyoggupta@users.noreply.github.com> Co-authored-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>	2025-10-25 15:26:45 -04:00
Wanli Jiang	95be56e56b	[TRTLLM-8238][feat] Add EVS support for nano-v2-vlm (#8024 ) Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com> Signed-off-by: Chang Liu <9713593+chang-l@users.noreply.github.com> Co-authored-by: Chang Liu <9713593+chang-l@users.noreply.github.com>	2025-10-25 05:43:27 -04:00
Simeng Liu	2b27810198	[https://nvbugs/5494718 ][fix] Fix Single GPU Multi-node issue and OOM on DGX Spark (#8514 ) Signed-off-by: Simeng Liu <simengl@nvidia.com>	2025-10-24 19:09:07 -07:00
jthomson04	02081e2390	[None][feat] Support KV Connector with Disagg Prefill Worker (#8246 ) Signed-off-by: jthomson04 <jwillthomson19@gmail.com>	2025-10-24 11:09:06 -07:00
Chang Liu	e47c787dd7	[TRTLLM-8535][feat] Support DeepSeek V3.2 with FP8 + BF16 KV cache/NVFP4 + BF16 KV cache (#8405 ) Signed-off-by: Fanrong Li <23290157+lfr-0531@users.noreply.github.com> Signed-off-by: Chang Liu <9713593+chang-l@users.noreply.github.com> Signed-off-by: Tracin <10434017+Tracin@users.noreply.github.com>	2025-10-24 13:40:41 -04:00
Yechan Kim	2d86d6be40	[TRTLLM-8737][feat] Support media_io_kwargs on trtllm-serve (#8528 ) Signed-off-by: yechank <161688079+yechank-nvidia@users.noreply.github.com>	2025-10-24 12:53:40 -04:00
Aurelien Chartier	cdf0403c64	[None][feat] Pass KvCacheRetentionConfig to torch LlmRequest (#8634 ) Signed-off-by: Aurelien Chartier <2567591+achartier@users.noreply.github.com>	2025-10-24 06:44:34 -07:00
Chuang Zhu	2420918e5b	[TRTLLM-7078][chore] optimal kvcache transfer for VWSA (#7952 ) Signed-off-by: Chuang Zhu <111838961+chuangz0@users.noreply.github.com>	2025-10-24 08:58:16 -04:00
Suyog Gupta	f512ddaeef	[None][feat] add skip condition in AutoDeploy's triton fused moe kernel (#8632 ) Signed-off-by: Suyog Gupta <41447211+suyoggupta@users.noreply.github.com>	2025-10-24 08:46:17 -04:00
QI JUN	6ee1c87595	[TRTLLM-8817][chore] Set default value of KvCacheConfig.free_gpu_memory_fraction explicitly (#8561 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>	2025-10-24 08:55:49 +08:00
h-guo18	23920223ab	[#4585 ][feat] Replace unified attention before export (#8303 ) Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>	2025-10-23 18:02:04 -04:00
QI JUN	cc81028547	[TRTLLM-8812][chore] Limit the scope of pybind based CacheTransceiverConfig (#8558 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>	2025-10-23 10:32:09 -04:00
Shijie	928247a3f9	[https://nvbugs/5451205 ][feat] Add cuBLASLt NVFP4 GEMM backend support (#7943 ) Signed-off-by: Shijie Wang <jaywan@nvidia.com>	2025-10-23 15:55:10 +08:00
Suyog Gupta	2956978da3	[None][feat] Enable rms norm fusion for Nemotron MOE (#8563 ) Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com> Signed-off-by: junq <22017000+QiJune@users.noreply.github.com> Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com> Signed-off-by: Suyog Gupta <41447211+suyoggupta@users.noreply.github.com> Co-authored-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com> Co-authored-by: QI JUN <22017000+QiJune@users.noreply.github.com> Co-authored-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-10-23 00:09:42 -04:00
sunnyqgg	ea3e0eea51	[TRTLLM-7954][feat] Target model KV cache rellocation (#8421 ) Signed-off-by: qgai <qgai@nvidia.com>	2025-10-23 09:36:50 +08:00
Anthony Chang	8a3b870e09	[None][feat] Update TRTLLM MoE MxFP4 cubins; autotune tileN (#8156 ) Signed-off-by: Anthony Chang <27950904+rosenrodt@users.noreply.github.com>	2025-10-23 09:14:18 +08:00
Anish Shanbhag	15de45d782	[TRTLLM-8682][chore] Remove auto_parallel module (#8329 ) Signed-off-by: Anish Shanbhag <ashanbhag@nvidia.com>	2025-10-22 20:53:08 -04:00
Leslie Fang	e5865de518	[TRTLLM-8754][chore] Refine PyTorchModelEngine with llm args (#8493 ) Signed-off-by: leslie-fang25 <leslief@nvidia.com>	2025-10-22 20:03:18 -04:00
Patrice Castonguay	879039f6d5	[https://nvbugs/5429636 ][feat] Kv transfer timeout (#8459 ) Signed-off-by: raayandhar <raayan.dhar@gmail.com> Signed-off-by: Patrice Castonguay <55748270+pcastonguay@users.noreply.github.com> Co-authored-by: raayandhar <raayan.dhar@gmail.com>	2025-10-22 09:29:02 -04:00
Yan Chunwei	f81caf5491	[None][chore] replace print_colored_debug with logger_debug (#8417 ) Signed-off-by: Yan Chunwei <328693+Superjomn@users.noreply.github.com>	2025-10-22 17:54:38 +08:00
sunnyqgg	90080e0e09	[https://nvbugs/5556020 ][fix] test_disaggregated_serving.py::TestLlama3_1_8BInstruct::test_eagle3 dimension mismatch (#8517 ) Signed-off-by: qgai <qgai@nvidia.com>	2025-10-22 09:58:22 +08:00
Leslie Fang	50d4e5bc06	[TRTLLM-8483][chore] Refine scheduler_config and peft_cache_config in create_py_executor (#8451 ) Signed-off-by: leslie-fang25 <leslief@nvidia.com>	2025-10-22 08:33:48 +08:00
Chenghao Zhang	bac9e8c2ad	[None][feat] AutoDeploy: Add Nemotron MOE support for AutoDeploy (#8469 )	2025-10-21 15:32:01 -07:00
Lucas Liebenwein	9b54b3bfaf	[None][chore] AutoDeploy: replace HF's deprecated keyword torch_dtype --> dtype (#8510 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-10-21 17:07:06 -04:00
YueWeng	8dc4aac5b6	[TRTLLM-8160][feat] Add max_total_draft_tokens (#8366 ) Signed-off-by: Yue Weng <25103990+yweng0828@users.noreply.github.com>	2025-10-21 11:11:04 -04:00
Bo Li	ebb62e17d8	[None][feat] Add alltoall to trtllm-gen MoE backend. (#8481 ) Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>	2025-10-21 12:42:54 +08:00
mpikulski	87eb5086fb	[None][fix] restore list[list[list[int]]] in add_token (#8502 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-10-20 22:34:57 -04:00
Yechan Kim	85d5aa7763	[None][feat] Support kv_cahce_reuse for HyperCLOVAX-Vision model (#7789 ) Signed-off-by: yechank <161688079+yechank-nvidia@users.noreply.github.com>	2025-10-21 11:11:24 +09:00
Suyog Gupta	7050b1ea49	[#8272 ][feat] Enable chunked prefill for SSMs in AutoDeploy (#8477 ) Signed-off-by: Suyog Gupta <41447211+suyoggupta@users.noreply.github.com>	2025-10-20 15:31:52 -07:00
Lucas Liebenwein	55c468b218	[#8461 ][feat] AutoDeploy: trtllm-serve bug fix + unit test (#8462 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-10-20 16:06:39 -04:00
Pamela Peng	b818a912d7	[https://nvbugs/5540752 ][fix] Support quantized Phi4 MM models (#8190 ) Signed-off-by: Pamela <179191831+pamelap-nvidia@users.noreply.github.com>	2025-10-20 06:36:09 -04:00
mpikulski	97ce0ecefe	[TRTLLM-8436][feat] batched sampling and top-k logprobs improvements (#8398 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-10-20 11:15:41 +02:00
ChristinaZ	c8b9998acb	[TRTLLM-8637][feat] Optimize the routing kernel for DeepseekV3 (MoE CUTLASS backend); Add support for KimiK2 and Qwen-next (MoE TRTLLM backend) (#7761 ) Signed-off-by: Christina Zhang <83400082+ChristinaZ@users.noreply.github.com>	2025-10-20 10:08:31 +08:00
Bo Deng	dd25595ae8	[TRTLLM-7964][infra] Set nixl to default cache transceiver backend (#7926 ) Signed-off-by: Bo Deng <deemod@nvidia.com>	2025-10-19 19:24:43 +08:00
jthomson04	852316886e	[None][fix] Fix KV event consumption (#6346 ) Signed-off-by: jthomson04 <jwillthomson19@gmail.com>	2025-10-18 15:41:26 -07:00
Lucas Liebenwein	41169fb20c	[None][feat] AutoDeploy: chunked prefill support (#8158 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-10-18 00:47:35 -07:00
QI JUN	4a8ac8dd62	[TRTLLM-8480][chore] clean create_py_executor API (#8412 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>	2025-10-17 23:52:02 -04:00
Wanli Jiang	58b43a6dab	[None][fix] Fix get_num_tokens_per_image for nano-v2-vlm (#8425 ) Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>	2025-10-18 08:51:35 +08:00
Kyle McGill	136e0e6882	[None][feat] Enable CUDA graph support for KvConnectorWorker API (#8275 ) Signed-off-by: Kyle McGill <kmcgill@nvidia.com> Signed-off-by: Kyle McGill <101670481+nv-kmcgill53@users.noreply.github.com>	2025-10-17 18:09:03 -04:00
h-guo18	55fed1873c	[None][chore] AutoDeploy: cleanup old inference optimizer configs (#8039 ) Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com> Co-authored-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-10-17 15:55:57 -04:00
Grzegorz Kwasniewski	bb7fdcebf4	[TRTLLM-8201][feat] Topological graph helpers (#8457 ) Signed-off-by: greg-kwasniewski1 <213329731+greg-kwasniewski1@users.noreply.github.com>	2025-10-17 12:34:19 -04:00
Tracin	dd06612d0e	[https://nvbugs/5540138 ][fix] Fix shape error when duplicating kv. (#8390 ) Signed-off-by: Tracin <10434017+Tracin@users.noreply.github.com>	2025-10-17 10:07:29 +08:00
John Calderon	46ee7acb33	[TRTLLM-6780][fix] Add multimodal data to dummy requests during memory profiling (#7539 ) Signed-off-by: John Calderon <johncalesp@gmail.com> Signed-off-by: John Calderon <jcalderon@nvidia.com> Signed-off-by: john calderon <jcalderon@nvidia.com> Signed-off-by: John Calderon <jcalderon@nvidia>	2025-10-16 17:49:22 +02:00
Jin Li	d594c2d0ff	[https://nvbugs/5537348 ][fix] Use device tensor index for MTP (#8062 ) Signed-off-by: Jin Li <59594262+liji-nv@users.noreply.github.com> Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com>	2025-10-16 22:46:19 +08:00
Yechan Kim	9587f099ac	[https://nvbugs/5547434 ][fix] Fix Qwen2.5-VL device_path error (#8057 ) Signed-off-by: yechank <161688079+yechank-nvidia@users.noreply.github.com> Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com>	2025-10-16 22:46:19 +08:00
Yukun He	179c7dc501	[https://nvbugs/5536131 ][fix] Fix illegal access issue when scale is not provided in Llama3/4. (#7960 ) Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com> Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com>	2025-10-16 22:46:19 +08:00
Enwei Zhu	57a4ef870a	[None][fix] Fix chunked prefill state of draft request (#8067 ) Signed-off-by: Enwei Zhu <21126786+syuoni@users.noreply.github.com> Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com>	2025-10-16 22:46:19 +08:00
Wangjue Yao	9865d3d770	[None][feat] Support cached tokens for Openai server (#7637 ) Signed-off-by: wjueyao <wyao123@terpmail.umd.edu> Co-authored-by: Pengyun Lin <81065165+LinPoly@users.noreply.github.com>	2025-10-16 20:51:37 +08:00
chinamaoge	ee588a73ac	[None][fix] Fix the error where checkpoint_dir is assigned as NONE wh… (#8401 ) Signed-off-by: maoge <maoge23@qq.com> Co-authored-by: maoge <maoge23@qq.com>	2025-10-16 13:37:43 +08:00
Min Yu	0a0159fdd8	[https://nvbugs/5378031 ] [feat] W4A8 AWQ MoE supports Per Expert Pre-quant Scale Factor for PyT backend (#7286 ) Signed-off-by: Min Yu <171526537+yumin066@users.noreply.github.com>	2025-10-16 11:07:48 +08:00
Wanli Jiang	ebf0e51206	[TRTLLM-8579][feat] Support quantized model for nano-v2-vlm (#8304 ) Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>	2025-10-16 09:44:11 +08:00
Chuang Zhu	40d129a415	[None][fix] Fix cache buffer size for window (#8320 ) Signed-off-by: Chuang Zhu <111838961+chuangz0@users.noreply.github.com>	2025-10-16 09:01:11 +08:00
HuiGao-NV	e265eb5fe9	[None][feat] reuse cudagraph memory pool in normal forward flow (#8095 ) Signed-off-by: Hui Gao <huig@nvidia.com>	2025-10-16 07:08:44 +08:00
dongfengy	7a0aa64973	[None][fix] Refactor triton paddings (#6980 ) Signed-off-by: Dongfeng Yu <dongfengy@nvidia.com> Signed-off-by: dongfengy <99041270+dongfengy@users.noreply.github.com> Co-authored-by: hlu1 <14827759+hlu1@users.noreply.github.com>	2025-10-15 12:59:01 -07:00
QI JUN	65ec01b257	[TRTLLM-8532][chore] clean warmup method of ModelEngine (#8264 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>	2025-10-15 08:40:58 -07:00
Yukun He	56c20665a9	[TRTLLM-4501][feat] Add input tensor pre-hook function API for the tuning process. (#6924 ) Some tunable ops require a more realistic data distribution, for instance, a shape-associated tensor. Thus, a customizable pre-hook function can be declared in the tuning config to modify the input tensor before the tuning process. Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>	2025-10-15 21:18:11 +08:00
mpikulski	93a4b7f1b6	[None][chore] update torch_dtype -> dtype in 'transformers' (#8263 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-10-15 17:09:30 +09:00
QI JUN	616d1df7a0	[None][chore] set the default value of max_num_tokens explicitly (#8208 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>	2025-10-14 23:03:02 -07:00
sychen52	6a6124dcb5	[OMNIML-2336][feat] w4a8 nvfp4 fp8 exports scale factor properly (#8180 ) Signed-off-by: Shiyang Chen <shiychen@nvidia.com> Co-authored-by: Shiyang Chen <shiychen@omniml-a6.nvidia.com>	2025-10-15 13:41:27 +08:00
Tailing Yuan	8444a50d3a	[None][fix] Fix is_post_quant_all2all_supported for MNNVL (#8355 ) Signed-off-by: Tailing Yuan <yuantailing@gmail.com>	2025-10-14 11:49:21 -07:00
Fanrong Li	0d20a8fd61	[TRTLLM-8536][feat] Add the sparse attention framework and one use case--RocketKV support (#8086 ) Signed-off-by: Fanrong Li <23290157+lfr-0531@users.noreply.github.com> Signed-off-by: yuhangh <58161490+heyuhhh@users.noreply.github.com> Co-authored-by: yuhangh <58161490+heyuhhh@users.noreply.github.com>	2025-10-14 08:23:16 -07:00
Yuxian Qiu	3450fe9944	[None][fix] Fix dummy load format for key models. (#7993 ) Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com>	2025-10-14 11:18:39 +08:00
Aurelien Chartier	9bc055faf1	[None][fix] Disable DeepGEMM for Qwen3 MoE Attention layers (#8087 ) Signed-off-by: Aurelien Chartier <2567591+achartier@users.noreply.github.com>	2025-10-13 18:38:47 -07:00
Lucas Liebenwein	22aa4ac08c	[None][feat] AutoDeploy: VLMs with subgraphs + cudagraph/compile (#8203 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-10-13 17:34:09 -07:00
Zheyu Fu	bac665e650	[TRTLLM-7412][feat] Turn off spec decode when the rolling average acceptance length drops below threshold. (#7283 ) Signed-off-by: Zheyu Fu <zheyuf@NVIDIA.com>	2025-10-13 15:51:14 -07:00
Grzegorz Kwasniewski	ea4658197f	[TRTLLM-6342][feat] Factory TP sharding of quantized models (#8123 ) Signed-off-by: greg-kwasniewski1 <213329731+greg-kwasniewski1@users.noreply.github.com> Co-authored-by: Suyog Gupta <41447211+suyoggupta@users.noreply.github.com>	2025-10-13 14:04:46 -07:00
Yuxian Qiu	bd740c9ba6	[None][fix] Avoid unnecessary concat in attn_output_gate case. (#8094 ) Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com>	2025-10-13 12:59:40 -07:00
Robin Kobus	db8c63b9b1	[TRTLLM-4517] [feat] Additional model outputs (#7206 ) Signed-off-by: Robin Kobus <19427718+Funatiq@users.noreply.github.com>	2025-10-13 15:33:18 +02:00
Po-Han Huang (NVIDIA)	6fc6f70a68	[https://nvbugs/5441729 ][test] Fix test_modeling_llama_min_latency.py failures (#7478 ) Signed-off-by: Po-Han Huang <pohanh@nvidia.com>	2025-10-13 15:35:02 +08:00
Leslie Fang	8d1b068b1a	[TRTLLM-8477][chore] Replace KvCacheConfigCpp with KvCacheConfig inside PyExecutor (#8259 ) Signed-off-by: leslie-fang25 <leslief@nvidia.com>	2025-10-13 14:55:36 +08:00
DylanChen-NV	d6e315e9ff	[None][feat] Add torch compile support for cuda core GEMM OP (#8261 ) Signed-off-by: Dylan Chen <191843203+DylanChen-NV@users.noreply.github.com>	2025-10-12 20:57:17 -07:00
amitz-nv	fac47e2826	[https://nvbugs/5510879 ][fix] Fix pytorch & TRT-python flows fused LoRA adapter modules weight split with TP>1 (#8063 ) Signed-off-by: Amit Zuker <203509407+amitz-nv@users.noreply.github.com>	2025-10-12 12:29:52 -07:00
kris1025	a7ea544dbe	[TRTLLM-7384][feat] enable rejection sampling for CDL (#7731 ) Signed-off-by: linquanh <linquanh@nvidia.com>	2025-10-12 20:38:48 +08:00
Ziyi Xiong	efd4ffa03b	[https://nvbugs/5534705 ][fix] Skip unnecessary CUDA graph capture (#8050 ) Signed-off-by: ziyixiong-nv <219238287+ziyixiong-nv@users.noreply.github.com>	2025-10-11 13:26:55 +08:00
QI JUN	48c15d805c	[https://nvbugs/5558167 ][fix] update canceled_req_ids correctly for canceled requests (#8207 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>	2025-10-10 18:58:26 +08:00
HuiGao-NV	795a051765	[None][chore] Print log with time for starting to load safetensor weights (#8218 ) Signed-off-by: Hui Gao <huig@nvidia.com>	2025-10-10 13:54:54 +08:00
mpikulski	7b6803b6e9	[TRTLLM-7769][chore] document the role of 'd2t' (#8174 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-10-09 13:13:50 -04:00
dongfengy	9f2a3ae88c	[None][fix] Restrict tinygemm use to certain SMs (#8182 ) Signed-off-by: Dongfeng Yu <dongfengy@nvidia.com> Signed-off-by: dongfengy <99041270+dongfengy@users.noreply.github.com>	2025-10-08 17:55:57 -07:00
mpikulski	8298e93bd8	[TRTLLM-8414][chore] BREAKING CHANGE: refine sampling strategy selection (#8132 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-10-08 15:46:50 +02:00
Sergey Klevtsov	017583a949	[https://nvbugs/5488576 ][fix] Propagate disable_finalize_fusion config flag in WIDEEP MoE backend (#8141 ) Signed-off-by: Sergey Klevtsov <sklevtsov@nvidia.com>	2025-10-07 14:44:54 -07:00
Mike Iovine	7facac077b	[None][fix] Fix MTP illegal memory access (#8161 ) Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com>	2025-10-07 14:02:55 -04:00
Faraz	27a5091fcb	[None][feat] GPT-OSS Sm120/Sm121 Support (#7937 ) Signed-off-by: Perkz Zheng <67892460+PerkzZheng@users.noreply.github.com> Signed-off-by: list <58580514+farazkh80@users.noreply.github.com> Signed-off-by: Vincent Huang <vincenth@nvidia.com> Co-authored-by: Perkz Zheng <67892460+PerkzZheng@users.noreply.github.com> Co-authored-by: Vincent Huang <vincenth@nvidia.com>	2025-10-06 16:59:06 -04:00
Izzy Putterman	f2657c1ae9	[None][fix] Eagle: Attention DP (#7939 ) Signed-off-by: Izzy Putterman <iputterman@nvidia.com>	2025-10-06 16:52:35 -04:00
Frida Hou	f6654f26a4	[#5255 ][autodeploy] Update FuseAllreduceResidualRMSNorm to use pattern matcher utility; remove fuse_collective (#7545 ) Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>	2025-10-05 01:15:46 -07:00
Frida Hou	744246d316	[None][autodeploy] small refactors on attention matching (#8079 ) Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>	2025-10-03 22:00:27 -07:00
Jonas Yang CN	88ea2c4ee9	[TRTLLM-7349][feat] Adding new orchestrator type -- ray (#7520 ) Signed-off-by: Erin Ho <14718778+hchings@users.noreply.github.com> Co-authored-by: Yuan Tong <13075180+tongyuantongyu@users.noreply.github.com> Co-authored-by: Erin Ho <14718778+hchings@users.noreply.github.com>	2025-10-04 08:12:24 +08:00
Lucas Liebenwein	9d098e3142	[None][feat] AutoDeploy: graph/module inputs with kwargs instead of args (#8137 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-10-03 16:53:42 -07:00
Michal Guzek	38da871db3	[TRTLLM-6496][feat] Add LoRa Torch tests for the latest NIM model list (#6806 ) Signed-off-by: Michal Guzek <mguzek@nvidia.com>	2025-10-03 12:10:48 -07:00
Mike Iovine	ca8291133a	[None][fix] Fix MTP 2-model (#8115 ) Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com> Signed-off-by: Mike Iovine <miovine@nvidia.com>	2025-10-03 10:13:50 -07:00
Lucas Liebenwein	aaf2c3c2e5	[None][feat] AutoDeploy: compiler backends based on nn.Module (#8126 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-10-03 12:14:21 -04:00
Ziyi Xiong	7bc2d9e993	[https://nvbugs/5537878 ][fix] Reserve an extra slot for padded batch (#7998 ) Signed-off-by: ziyixiong-nv <219238287+ziyixiong-nv@users.noreply.github.com>	2025-10-03 08:42:52 -07:00
Suyog Gupta	d8215241d8	[None][feat] AutoDeploy add autotuning when capturing cudagraphs (#8120 ) Signed-off-by: Suyog Gupta <41447211+suyoggupta@users.noreply.github.com>	2025-10-03 08:33:21 -07:00
Aurelien Chartier	9db4366903	[None][fix] Fix Qwen3 FP8 per-tensor when requesting TRTLLM-GEN MoE backend (#8075 ) Signed-off-by: Aurelien Chartier <2567591+achartier@users.noreply.github.com>	2025-10-03 07:52:52 -07:00
Lucas Liebenwein	5faa5e9dd8	[None][feat] AutoDeploy: dive deeper into token generation bugs + enable_block_reuse (#8108 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-10-03 04:57:26 -07:00
Nikita Korobov	9b3d7cc3e6	[None][feat] Update TRT-LLM Gen MoE kernels (#7970 ) Signed-off-by: Nikita Korobov <14355239+nekorobov@users.noreply.github.com>	2025-10-03 09:22:45 +08:00
Yilin Fan	01423ac183	[None][feat] perf_metrics endpoint functionality improvement (#8005 ) Signed-off-by: Yilin Fan <206948969+nv-yilinf@users.noreply.github.com> Signed-off-by: nv-yilinf <206948969+nv-yilinf@users.noreply.github.com>	2025-10-02 17:43:25 -07:00
Grzegorz Kwasniewski	a5b59fd31d	[TRTLLM-6342][bug] Patched incorrect starcoder tp config (#8118 ) Signed-off-by: greg-kwasniewski1 <213329731+greg-kwasniewski1@users.noreply.github.com>	2025-10-02 18:41:59 -04:00
Daniel Cámpora	ab433b7228	[None][fix] Fix access to new tokens in sampler. (#7958 ) Signed-off-by: Daniel Campora <961215+dcampora@users.noreply.github.com>	2025-10-02 15:41:21 -04:00
Patrice Castonguay	fefa7d8fa3	[None][feat] Support for cancelling requests with disaggregation (#8114 ) Signed-off-by: Shunkang <182541032+Shunkangz@users.noreply.github.co> Signed-off-by: Patrice Castonguay <55748270+pcastonguay@users.noreply.github.com> Co-authored-by: Shunkang <182541032+Shunkangz@users.noreply.github.co>	2025-10-02 11:04:26 -07:00
dongfengy	6568e565db	[TRTLLM-7775][feat] Integrate tinygemm2 for gpt-oss (#7916 ) Signed-off-by: Dongfeng Yu <dongfengy@nvidia.com> Signed-off-by: dongfengy <99041270+dongfengy@users.noreply.github.com> Co-authored-by: Jin Li <59594262+liji-nv@users.noreply.github.com>	2025-10-02 10:47:04 -07:00
yifeizhang-c	34d158b6da	[TRTLLM-6589][feat] Support CUDA graph for DeepEP (#7514 ) Signed-off-by: Yifei Zhang <219273404+yifeizhang-c@users.noreply.github.com>	2025-10-02 10:13:24 -07:00
Chang Liu	726ac07cc0	[https://nvbugs/5549081 ][fix] Fix device id assignment for some vision models (#8070 ) Signed-off-by: Chang Liu (Enterprise Products) <9713593+chang-l@users.noreply.github.com> Signed-off-by: Chang Liu <9713593+chang-l@users.noreply.github.com>	2025-10-01 23:28:05 -04:00
brb-nv	bd3d0ad233	[TRTLLM-7733][feat] Executor changes to support helix parallelism (#7972 ) Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>	2025-10-01 22:13:03 -04:00
Izzy Putterman	1ad7bc4c78	[None][feat] Draft: Save state first pass (#7012 ) Signed-off-by: Izzy Putterman <iputterman@nvidia.com>	2025-10-01 18:40:55 -04:00
Frida Hou	de99e23696	[#5860 ][feat] Add ModelOPT INT4 awq fake quant support in AutoDeploy (#7770 ) Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>	2025-10-01 13:13:45 -07:00
Yibin Li	d7581bb551	[TRTLLM-8031][feat] Add chunked return_generation_logits logic (#7831 ) Signed-off-by: Yibin Li <109242046+yibinl-nvidia@users.noreply.github.com>	2025-10-01 12:47:07 -04:00
Grzegorz Kwasniewski	6fd225833c	[TRTLLM-6342][bug] Fix shape propagation after TP sharding (#7912 ) Signed-off-by: greg-kwasniewski1 <213329731+greg-kwasniewski1@users.noreply.github.com>	2025-10-01 11:15:46 -04:00
sychen52	ba8abeab10	[OMNIML-2336][feat] add W4A8 NVFP4 FP8 fused moe (#7968 ) Signed-off-by: Shiyang Chen <shiychen@nvidia.com>	2025-10-01 02:39:33 -04:00
peaceh-nv	808e556c79	[None][fix] : Fix OOM issue when dp padding is enabled (#8052 ) Signed-off-by: peaceh <103117813+peaceh-nv@users.noreply.github.com>	2025-10-01 09:10:00 +08:00
brb-nv	84aa3c981e	[None][chore] Waive failing MNNVL alltoall multi-gpu test (#8106 ) Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>	2025-09-30 20:05:42 -04:00
Guoming Zhang	b4be0d2e4c	[None][chore] Refine qwen3-next implementation. (#8064 ) Signed-off-by: nv-guomingz <137257613+nv-guomingz@users.noreply.github.com>	2025-09-30 15:05:13 -04:00
Yechan Kim	948b8b9569	[None][fix] Fix CUDA graph for Qwen2.5-VL (#8047 ) Signed-off-by: yechank <161688079+yechank-nvidia@users.noreply.github.com>	2025-09-30 14:40:03 +08:00
Kaiyu Xie	b0cb9ca50e	[None] [test] Add MNNVL AlltoAll tests to pre-merge (#7466 ) Signed-off-by: Kaiyu Xie <26294424+kaiyux@users.noreply.github.com> Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com> Co-authored-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com>	2025-09-29 23:12:24 -04:00
Lucas Liebenwein	dcfd3ef81c	[#4593 ][feat] AutoDeploy: Linear Attention Support (SSM + causal_conv + Bamba + Nemotron-H) (#8068 ) Signed-off-by: William Zhang <133824995+2ez4bz@users.noreply.github.com> Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com> Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com> Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com> Signed-off-by: Suyog Gupta <41447211+suyoggupta@users.noreply.github.com> Co-authored-by: William Zhang <133824995+2ez4bz@users.noreply.github.com> Co-authored-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com> Co-authored-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com> Co-authored-by: Suyog Gupta <41447211+suyoggupta@users.noreply.github.com>	2025-09-29 22:41:06 -04:00
Cao Dong	62010c0ab7	[None][feat] Return topk logprobs in torch backend (#7976 ) Signed-off-by: Cao Dong <87467313+dcaox@users.noreply.github.com>	2025-09-30 09:32:37 +08:00
Cheng Hang	cdce68c3e0	[TRTLLM-6741][fix] Add heuristics for lm head tp size when `enable_lm_head_tp_in_adp=True` (#7891 ) Signed-off-by: Cheng Hang <chang@nvidia.com> Co-authored-by: Yanchao Lu <yanchaol@nvidia.com>	2025-09-30 09:24:35 +08:00
bhsueh_NV	38d6e4e60b	[None][feat] Support Qwen3 next (#7892 ) Signed-off-by: mengw <12670782+wm2012011492@users.noreply.github.com> Signed-off-by: bhsueh <11360707+byshiue@users.noreply.github.com> Signed-off-by: nv-guomingz <137257613+nv-guomingz@users.noreply.github.com> Co-authored-by: nv-guomingz <137257613+nv-guomingz@users.noreply.github.com>	2025-09-29 21:16:07 +08:00
mpikulski	a0d489a8d5	[TRTLLM-7728][perf] improve batched sampling perf for contiguous batches (#7908 ) Signed-off-by: ixlmar <206748156+ixlmar@users.noreply.github.com>	2025-09-29 13:32:50 +01:00
Gal Hubara-Agam	b2095aa074	[#4674 ][bugfix] AutoDeploy Fix memory leak in fuse_moe (#7844 ) Delete the unstacked weights immediately to save GPU memory, cleanup occurs automatically after the transformation, but for large models we'll run out of memory during the transformation itself. Signed-off-by: Gal Hubara Agam <96368689+galagam@users.noreply.github.com>	2025-09-29 11:01:07 +03:00
Void	7f1e2dba92	[None][fix] only support deepep post quant all2all on nvfp4 (#8041 ) Signed-off-by: Yilin Zhang <18275976+yilin-void@users.noreply.github.com>	2025-09-29 14:37:50 +08:00
Tailing Yuan	985b79ca82	[TRTLLM-8348][feat] Speed up concat k and copy k_nope in context phase using torch.compile (#8044 ) Signed-off-by: Tailing Yuan <yuantailing@gmail.com>	2025-09-29 13:28:12 +08:00
Eran Geva	9cea6bfb30	[#7288 ][feat] Added AutoDeploy backend support to test_perf.py (#7588 ) Signed-off-by: Eran Geva <19514940+MrGeva@users.noreply.github.com>	2025-09-28 21:21:27 -07:00
Zongfei Jing	e9f26feeb6	[None][chore] Cherry-pick from (#7598 ) Make low_precision_combine as a llm arg (#7898 ) Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com>	2025-09-28 22:32:33 -04:00
Yukun He	28b9a81c58	[TRTLLM-4500][feat] Add serialization/deserialization options for AutoTuner profiling cache (#7738 ) To achieve determinism for the AutoTuner profiling cache, serialization and deserialization are introduced to store the cache on disk in JSON format. Use TLLM_AUTOTUNER_CACHE_PATH to indicate the path where the cache file should be stored: Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>	2025-09-29 07:40:51 +08:00
Guoming Zhang	3ba4bf6e70	[None][chore] Disable concurrent weights loading for _load_weights_im… (#8034 ) Signed-off-by: nv-guomingz <137257613+nv-guomingz@users.noreply.github.com>	2025-09-28 07:11:16 -04:00
ChristinaZ	95eac2cda7	[https://nvbugs/5537738 ][fix] Add fp8 post-quant allgather support (#8008 ) Signed-off-by: Christina Zhang <83400082+ChristinaZ@users.noreply.github.com>	2025-09-28 15:32:45 +08:00
Aurelien Chartier	77b68d9d7d	[https://nvbugs/5461712 ] [fix] Use DG for Qwen3 Linear layers (#8030 ) Signed-off-by: Aurelien Chartier <2567591+achartier@users.noreply.github.com>	2025-09-28 10:33:36 +08:00
Xianjie Qiao	c8f98b3065	[None] [feat] Update disagg gen-only benchmark. (#7917 ) Signed-off-by: Xianjie <5410381+qiaoxj07@users.noreply.github.com>	2025-09-28 09:56:56 +08:00
Iman Tabrizian	33282351a2	[TRTLLM-6106][feat] Add support for KVCache transfer from KVCache reuse path (#6348 ) Signed-off-by: Iman Tabrizian <10105175+tabrizian@users.noreply.github.com>	2025-09-27 19:29:30 -04:00
Frida Hou	a36b48bcab	[#5860 ][autodeploy] GPT-OSS MXFP4 support (#7451 ) Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>	2025-09-26 15:36:06 -07:00
Jhao-Ting Chen	c33f43e13a	[https://nvbugs/5518713 ][fix] Trtllm-gen moe backend for blockwise fp8 ckpt (Qwen3-235B-A22B-FP8) (#7856 ) Signed-off-by: Jhao-Ting Chen <jhaotingc@nvidia.com>	2025-09-26 14:29:32 -07:00
Mike Iovine	d7087015f1	[TRTLLM-8271][fix] Fix CDL overlap scheduling performance (#7971 ) Signed-off-by: Mike Iovine <6158008+mikeiovine@users.noreply.github.com>	2025-09-26 16:05:10 -04:00
YueWeng	a4243f0da5	[TRTLLM-6393][feat] add static tree sampling and verification (#7161 ) Signed-off-by: Yue Weng <25103990+yweng0828@users.noreply.github.com>	2025-09-26 13:16:16 -04:00
HuiGao-NV	f4d3be4bbc	[None][feat] Add a standalone buffer cache class and reuse buffers between cduagraph and no-graph flow (#7669 ) Signed-off-by: Hui Gao <huig@nvidia.com>	2025-09-26 07:28:06 -07:00
HuiGao-NV	a9965d84e0	[None][chore] Report NCCL error message but not OOM when NCCL error happens (#8009 ) Signed-off-by: Hui Gao <huig@nvidia.com>	2025-09-25 23:07:32 -07:00
peaceh-nv	55ce70060e	[https://nvbugs/5451740 ][fix] Add DP padding back on SM120 (#7965 ) Signed-off-by: peaceh <103117813+peaceh-nv@users.noreply.github.com>	2025-09-26 13:59:54 +08:00
Lucas Liebenwein	3a96d75a3c	[https://nvbugs/5527956 ][fix] AutoDeploy: fix IMA due to outdated metadata (#8002 ) Signed-off-by: Lucas Liebenwein <11156568+lucaslie@users.noreply.github.com>	2025-09-25 22:05:55 -07:00
sunnyqgg	2e5850c28a	[TRTLLM-7330][feat] Eagle3 cuda graph support for the first draft model inference (#7363 ) Signed-off-by: qgai <qgai@nvidia.com>	2025-09-26 11:28:05 +08:00
dongfengy	1eb653146a	[https://nvbugs/5525951 ][fix] Clarify that PP is not supported for GPTOSS (#7911 ) Signed-off-by: Dongfeng Yu <dongfengy@nvidia.com>	2025-09-25 12:54:18 -07:00
QI JUN	1529a6f22d	[None][chore] extract weights loading related logic to model loader (#7579 ) Signed-off-by: junq <22017000+QiJune@users.noreply.github.com>	2025-09-25 10:19:22 -07:00
xxi	57ff5f4c0d	[None][fix] fix a bug in wideEp use DeepEP with num_chunks > 1 (#7954 ) Signed-off-by: xxi <xxi@nvidia.com>	2025-09-25 07:53:42 -07:00
Matthias Jouanneaux	eda1467061	[TRTLLM-5966][feat] Helix: add alltoall op (#6815 ) Signed-off-by: Matthias Jouanneaux <mjoux@nvidia.com>	2025-09-25 07:18:29 -07:00
Guoming Zhang	202bed4574	[None][chroe] Rename TensorRT-LLM to TensorRT LLM for source code. (#7851 ) Signed-off-by: nv-guomingz <137257613+nv-guomingz@users.noreply.github.com> Signed-off-by: Wangshanshan <30051912+dominicshanshan@users.noreply.github.com>	2025-09-25 21:02:35 +08:00

... 4 5 6 7 8 ...

1428 Commits