TensorRT-LLMs/cpp/tests/unit_tests/kernels/allReduce
Zongfei Jing c7548ad72c
perf: Add optimizations for deepseek in min latency mode (#3093)
* Add optimizations for deepseek min latency

Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com>

* Fix compile error

Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com>

* Update internal cutlass kernel libs

Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com>

* Format code

Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com>

* Resolve conflicts

Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com>

---------

Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com>
2025-04-02 09:05:24 +08:00
..
allReduceFusionTest.cu perf: Add optimizations for deepseek in min latency mode (#3093) 2025-04-02 09:05:24 +08:00
allReduceKernelTest.cu Update TensorRT-LLM (#2792) 2025-02-18 21:27:39 +08:00
gemmAllReduceTest.cu Update (#2978) 2025-03-23 16:39:35 +08:00