TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

History

Maurits de Groot 2d0c9b383f [None][fix] Updated blog9_Deploying_GPT_OSS_on_TRTLLM (#7260 ) Signed-off-by: Maurits de Groot <63357890+Maurits-de-Groot@users.noreply.github.com>		2025-08-26 11:26:19 -04:00
..
media	[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )	2025-08-01 16:46:15 +08:00
tech_blog	[None][fix] Updated blog9_Deploying_GPT_OSS_on_TRTLLM (#7260 )	2025-08-26 11:26:19 -04:00
Best_perf_practice_on_DeepSeek-R1_in_TensorRT-LLM.md	[None][doc] Modify the description for mla chunked context (#6929 )	2025-08-15 12:52:26 +08:00
Falcon180B-H200.md	[https://nvbugs/5423962 ][fix] Address broken links (#6531 )	2025-08-07 16:00:05 -04:00
H100vsA100.md	chore: Mass integration of release/0.20 (#4898 )	2025-06-08 23:26:26 +08:00
H200launch.md	chore: Mass integration of release/0.20 (#4898 )	2025-06-08 23:26:26 +08:00
quantization-in-TRT-LLM.md	Update TensorRT-LLM (#2792 )	2025-02-18 21:27:39 +08:00
XQA-kernel.md	Update TensorRT-LLM (#2008 )	2024-07-23 23:05:09 +08:00