TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

History

Kefeng-Duan f5b6d453aa doc： DS r1 min latency blog (#4386 ) * add best perf practice on DSR1 Signed-off-by: Jun Yang <143764042+juney-nvidia@users.noreply.github.com> * add ds-r1 min latency tech blog Signed-off-by: Jun Yang <143764042+juney-nvidia@users.noreply.github.com> * rm redundant doc Signed-off-by: Jun Yang <143764042+juney-nvidia@users.noreply.github.com> * refine table content Signed-off-by: Jun Yang <143764042+juney-nvidia@users.noreply.github.com> * refine table content Signed-off-by: Jun Yang <143764042+juney-nvidia@users.noreply.github.com> * relative path for images Signed-off-by: Jun Yang <143764042+juney-nvidia@users.noreply.github.com> * refine precommit Signed-off-by: Jun Yang <143764042+juney-nvidia@users.noreply.github.com> * pr4280 is merged Signed-off-by: Jun Yang <143764042+juney-nvidia@users.noreply.github.com> --------- Signed-off-by: Jun Yang <143764042+juney-nvidia@users.noreply.github.com>		2025-05-16 20:20:28 +08:00
..
_static	chore: Mass Integration 0.19 (#4255 )	2025-05-16 10:53:25 +02:00
_templates	Update TensorRT-LLM (#1725 )	2024-06-04 20:26:32 +08:00
advanced	Revert "feat: Low Precision Allreduce for PCIe based GPU" (#4340 )	2025-05-15 09:52:39 +08:00
architecture	chore: Mass Integration 0.19 (#4255 )	2025-05-16 10:53:25 +02:00
blogs	doc： DS r1 min latency blog (#4386 )	2025-05-16 20:20:28 +08:00
commands	feat: enhance trtllm serve multimodal (#3757 )	2025-05-15 16:16:31 -07:00
dev-on-cloud	doc: add doc ahout developent on cloud or runpod (#3194 )	2025-04-02 18:10:56 +08:00
examples	doc: refactor trtllm-serve examples and doc (#3187 )	2025-04-04 11:40:43 +08:00
installation	[TRTLLM-5171] chore: Remove GptSession/V1 from TRT workflow (#4092 )	2025-05-14 23:10:04 +02:00
llm-api	doc: fix path after examples migration (#3814 )	2025-04-24 02:36:45 +08:00
media	L4 added to readme (#3301 )	2025-04-06 19:09:28 +08:00
performance	chore: Mass Integration 0.19 (#4255 )	2025-05-16 10:53:25 +02:00
python-api	Update TensorRT-LLM (#1492 )	2024-04-24 14:44:22 +08:00
reference	chore: Mass Integration 0.19 (#4255 )	2025-05-16 10:53:25 +02:00
torch	Remove dummy forward path (#3669 )	2025-04-18 16:17:50 +08:00
conf.py	chore: Mass Integration 0.19 (#4255 )	2025-05-16 10:53:25 +02:00
helper.py	doc: refactor trtllm-serve examples and doc (#3187 )	2025-04-04 11:40:43 +08:00
index.rst	doc： DS r1 min latency blog (#4386 )	2025-05-16 20:20:28 +08:00
key-features.md	Update TensorRT-LLM (#2562 )	2024-12-11 00:31:05 -08:00
overview.md	chore: Mass integration of release/0.18 (#3421 )	2025-04-16 10:03:29 +08:00
quick-start-guide.md	chore: Mass Integration 0.19 (#4255 )	2025-05-16 10:53:25 +02:00
release-notes.md	chore: Mass Integration 0.19 (#4255 )	2025-05-16 10:53:25 +02:00
torch.md	chore: Mass Integration 0.19 (#4255 )	2025-05-16 10:53:25 +02:00