TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

History

Yukun He 93a0fd0a23 [TRTLLM-6445] feat: Enable AllReduce-associated fusion patterns in Llama3/4. (#6205 ) Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>		2025-07-28 09:36:26 +08:00
..
allReduceFusionKernels.cu	[TRTLLM-6445] feat: Enable AllReduce-associated fusion patterns in Llama3/4. (#6205 )	2025-07-28 09:36:26 +08:00
allReduceFusionKernels.h	Feat/ds r1 min latency opt round3, add router gemm, fused a gemm, PDL (#4560 )	2025-06-14 17:36:22 +08:00
allReduceWorkspace.cu
allReduceWorkspace.h
customLowPrecisionAllReduceKernels.cu
customLowPrecisionAllReduceKernels.h
mnnvlTwoShotAllreduceKernels.cu	[fix][nvbugs/5399355] Fix Lamport buffer clear issue for MNNVL TwoShot Allreduce and add FP16 support. (#6237 )	2025-07-25 08:01:40 +08:00
mnnvlTwoShotAllreduceKernels.h
moeAllReduceFusionKernels.cu	Feat/ds r1 min latency opt round3, add router gemm, fused a gemm, PDL (#4560 )	2025-06-14 17:36:22 +08:00
moeAllReduceFusionKernels.h