TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-22 03:35:00 +08:00

History

Zongfei Jing c7548ad72c perf: Add optimizations for deepseek in min latency mode (#3093 ) * Add optimizations for deepseek min latency Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com> * Fix compile error Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com> * Update internal cutlass kernel libs Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com> * Format code Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com> * Resolve conflicts Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com> --------- Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com>		2025-04-02 09:05:24 +08:00
..
test_ar_residual_norm.py	test: reorganize tests folder hierarchy (#2996 )	2025-03-27 12:07:53 +08:00
test_deepseek_allreduce.py	perf: Add optimizations for deepseek in min latency mode (#3093 )	2025-04-02 09:05:24 +08:00
test_embedding.py	Update TensorRT-LLM (#2936 )	2025-03-18 21:25:19 +08:00
test_linear.py	Update TensorRT-LLM (#2936 )	2025-03-18 21:25:19 +08:00
test_star_attention_input.jsonl	Update TensorRT-LLM (#2936 )	2025-03-18 21:25:19 +08:00
test_star_attention.py	move BuildConfig functional args to llmargs (#3036 )	2025-03-29 02:20:18 +08:00
test_user_buffers.py	None - Add one-shot version for UB AR NORM FP16/BF16 (#2995 )	2025-03-31 11:16:03 +08:00