TensorRT-LLMs/cpp/tensorrt_llm/kernels/communicationKernels
Yukun He 93a0fd0a23
[TRTLLM-6445] feat: Enable AllReduce-associated fusion patterns in Llama3/4. (#6205)
Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>
2025-07-28 09:36:26 +08:00
..
allReduceFusionKernels.cu [TRTLLM-6445] feat: Enable AllReduce-associated fusion patterns in Llama3/4. (#6205) 2025-07-28 09:36:26 +08:00
allReduceFusionKernels.h Feat/ds r1 min latency opt round3, add router gemm, fused a gemm, PDL (#4560) 2025-06-14 17:36:22 +08:00
allReduceWorkspace.cu
allReduceWorkspace.h
customLowPrecisionAllReduceKernels.cu
customLowPrecisionAllReduceKernels.h
mnnvlTwoShotAllreduceKernels.cu [fix][nvbugs/5399355] Fix Lamport buffer clear issue for MNNVL TwoShot Allreduce and add FP16 support. (#6237) 2025-07-25 08:01:40 +08:00
mnnvlTwoShotAllreduceKernels.h
moeAllReduceFusionKernels.cu Feat/ds r1 min latency opt round3, add router gemm, fused a gemm, PDL (#4560) 2025-06-14 17:36:22 +08:00
moeAllReduceFusionKernels.h