TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-22 11:42:41 +08:00

History

Perkz Zheng 4d711be8f4 Feat: add sliding-window-attention generation-phase kernels on Blackwell (#4564 ) * move cubins to LFS Signed-off-by: Perkz Zheng <67892460+PerkzZheng@users.noreply.github.com> * update cubins Signed-off-by: Perkz Zheng <67892460+PerkzZheng@users.noreply.github.com> * add sliding-window-attention generation-phase kernels on Blackwell Signed-off-by: Perkz Zheng <67892460+PerkzZheng@users.noreply.github.com> * address comments Signed-off-by: Perkz Zheng <67892460+PerkzZheng@users.noreply.github.com> --------- Signed-off-by: Perkz Zheng <67892460+PerkzZheng@users.noreply.github.com>		2025-05-26 09:06:33 +08:00
..
trtllmGen_bmm_export	Feat: add sliding-window-attention generation-phase kernels on Blackwell (#4564 )	2025-05-26 09:06:33 +08:00
CMakeLists.txt	feat: TRT-LLM Gen integration for BMM and MoE refactoring (#4280 )	2025-05-16 13:31:53 +02:00
KernelRunner.cpp	feat: TRT-LLM Gen integration for BMM and MoE refactoring (#4280 )	2025-05-16 13:31:53 +02:00
KernelRunner.h	feat: TRT-LLM Gen integration for BMM and MoE refactoring (#4280 )	2025-05-16 13:31:53 +02:00