TensorRT-LLMs/cpp/tensorrt_llm/plugins/mixtureOfExperts
Jinyang Yuan 0a0f93d4a8
[None][fix] Fix the performance issue of FP8 blockwise grouped GEMM when using attention DP (#8501)
Signed-off-by: Jinyang Yuan <154768711+jinyangyuan-nvidia@users.noreply.github.com>
2025-10-27 10:18:19 +08:00
..
CMakeLists.txt Update TensorRT-LLM (#524) 2023-12-01 22:27:51 +08:00
mixtureOfExpertsPlugin.cpp [None][fix] Fix the performance issue of FP8 blockwise grouped GEMM when using attention DP (#8501) 2025-10-27 10:18:19 +08:00
mixtureOfExpertsPlugin.h [None][perf] Make finalize fusion part of the tactic selection logic (#6915) 2025-08-21 14:08:03 -07:00