Xiwen Yu
|
5bd50d477e
|
update mha cubins and support 103a
Signed-off-by: Xiwen Yu <xiweny@nvidia.com>
|
2025-09-02 19:26:24 -07:00 |
|
Xiwen Yu
|
38ef850552
|
Merge remote-tracking branch 'gitlab/main' into user/xiweny/merge_0901
Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com>
|
2025-09-01 11:46:44 +08:00 |
|
Tian Zheng
|
e257cb3533
|
[None][feat] Support NVFP4 KV Cache (#6244)
Signed-off-by: Tian Zheng <29906817+Tom-Zheng@users.noreply.github.com>
|
2025-09-01 09:24:52 +08:00 |
|
Xiwen Yu
|
fa8b52ed33
|
fix more sm version check
Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com>
|
2025-08-23 15:17:59 +08:00 |
|
Xiwen Yu
|
4a95d88ce2
|
revert tlg kernels for ease of merge
Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com>
|
2025-08-19 11:44:36 +08:00 |
|
Tian Zheng
|
3a94d80839
|
Update SM100f cubins
Signed-off-by: Tian Zheng <29906817+Tom-Zheng@users.noreply.github.com>
|
2025-08-06 14:25:00 +08:00 |
|
zhhuang-nv
|
94e6167879
|
optimize cudaMemGetInfo for TllmGenFmhaRunner (#3907)
Signed-off-by: Zhen Huang <145532724+zhhuang-nv@users.noreply.github.com>
|
2025-04-29 14:17:07 +08:00 |
|
Dan Blanaru
|
16d2467ea8
|
Update TensorRT-LLM (#2755)
* Update TensorRT-LLM
---------
Co-authored-by: Denis Kayshev <topenkoff@gmail.com>
Co-authored-by: akhoroshev <arthoroshev@gmail.com>
Co-authored-by: Patrick Reiter Horn <patrick.horn@gmail.com>
Update
|
2025-02-11 03:01:00 +00:00 |
|