TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-24 04:33:04 +08:00

History

Neta Zmora 34dc6869f3 [#8732 ][feat] Update TRTLLM Cutlass MoE kernels with ReLU2 (#9011 ) Update TRTLLM Cutlass MoE kernels with ReLU2 activation. Nemotron-6 requires ReLU2 (i.e. squared ReLU) MoE activation function. The PR adds this and adds an API to set the activation function, in general. The ReLU2 changes are based on this FlashInfer PR: https://github.com/flashinfer-ai/flashinfer/pull/1954. The PR also updates the Auto Deploy MoE backend for 16-bit and FP8 from Triton (`torch.ops.auto_deploy.triton_moe_fused`, `torch.ops.auto_deploy.triton_quant_fp8_moe`) to TRTLLM/Cutlass (`torch.ops.auto_deploy.trtllm_moe_fused`, `torch.ops.auto_deploy.trtllm_quant_fp8_moe_fused`). Signed-off-by: Neta Zmora <96238833+nzmora-nvidia@users.noreply.github.com> Signed-off-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com> Co-authored-by: Chenghao Zhang <211069071+nvchenghaoz@users.noreply.github.com>		2025-11-13 16:54:45 -08:00
..
__init__.py	[None][feat] Support Qwen3 next (#7892 )	2025-09-29 21:16:07 +08:00
cpp_custom_ops.py	[None][feat] Add customized topk and related unit tests for DSA (#8882 )	2025-11-10 03:35:35 -08:00
cute_dsl_custom_ops.py	[None][fix] Fix cute dsl nvfp4 gemm autotune issue (#8761 )	2025-11-03 22:55:45 +08:00
flashinfer_custom_ops.py	[None][feat] Support Qwen3 next (#7892 )	2025-09-29 21:16:07 +08:00
torch_custom_ops.py	[#8732 ][feat] Update TRTLLM Cutlass MoE kernels with ReLU2 (#9011 )	2025-11-13 16:54:45 -08:00
trtllm_gen_custom_ops.py	[None][fix] support topk autotuner input for expert slot per group larger than 32 (#9087 )	2025-11-14 08:37:20 +08:00
userbuffers_custom_ops.py	feat: Introduce UB allocator for pytorch flow (#3257 )	2025-04-08 18:39:49 +08:00