TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-23 20:23:08 +08:00

History

Anthony Chang 4742c130db [None][feat] Improve TRTLLM MoE in small hidden size throughput cases (#9377 ) Signed-off-by: Anthony Chang <27950904+rosenrodt@users.noreply.github.com>		2025-11-25 09:09:27 +01:00
..
trtllmGen_bmm_export	[None][feat] Improve TRTLLM MoE in small hidden size throughput cases (#9377 )	2025-11-25 09:09:27 +01:00
CMakeLists.txt	feat: TRT-LLM Gen integration for BMM and MoE refactoring (#4280 )	2025-05-16 13:31:53 +02:00
KernelRunner.cpp	[None][feat] Update TRTLLM MoE cubins; reduce mxfp4 weight padding requirement; tighten TMA bound (#9025 )	2025-11-17 10:04:29 +08:00
KernelRunner.h	[None][feat] Update TRTLLM MoE cubins; reduce mxfp4 weight padding requirement; tighten TMA bound (#9025 )	2025-11-17 10:04:29 +08:00