TensorRT-LLMs/docs/source
Zongfei Jing 7eee9a9d28
doc: Update doc for Deepseek min latency (#3717)
* Tidy code

Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com>

* Update doc for min latency deepseek

Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com>

* Throw exception for RouterKernel when not running on sm90+

Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com>

---------

Signed-off-by: Zongfei Jing <20381269+zongfeijing@users.noreply.github.com>
2025-04-22 23:07:59 +08:00
..
_templates Update TensorRT-LLM (#1725) 2024-06-04 20:26:32 +08:00
advanced disagg perf tune doc (#3516) 2025-04-14 16:05:06 +08:00
architecture Update TensorRT-LLM (#2562) 2024-12-11 00:31:05 -08:00
blogs doc: Update doc for Deepseek min latency (#3717) 2025-04-22 23:07:59 +08:00
commands feat: trtllm-serve multimodal support (#3590) 2025-04-19 05:01:28 +08:00
dev-on-cloud doc: add doc ahout developent on cloud or runpod (#3194) 2025-04-02 18:10:56 +08:00
examples doc: refactor trtllm-serve examples and doc (#3187) 2025-04-04 11:40:43 +08:00
installation relax the limitation of setuptools (#2992) 2025-03-24 13:36:10 +08:00
llm-api Update TensorRT-LLM (#2755) 2025-02-11 03:01:00 +00:00
media L4 added to readme (#3301) 2025-04-06 19:09:28 +08:00
performance feat: adding multimodal (only image for now) support in trtllm-bench (#3490) 2025-04-18 07:06:16 +08:00
python-api Update TensorRT-LLM (#1492) 2024-04-24 14:44:22 +08:00
reference chore: Mass integration of release/0.18 (#3421) 2025-04-16 10:03:29 +08:00
torch Remove dummy forward path (#3669) 2025-04-18 16:17:50 +08:00
conf.py doc: add genai-perf benchmark & slurm multi-node for trtllm-serve doc (#3407) 2025-04-16 00:11:58 +08:00
helper.py doc: refactor trtllm-serve examples and doc (#3187) 2025-04-04 11:40:43 +08:00
index.rst doc: refactor trtllm-serve examples and doc (#3187) 2025-04-04 11:40:43 +08:00
key-features.md Update TensorRT-LLM (#2562) 2024-12-11 00:31:05 -08:00
overview.md chore: Mass integration of release/0.18 (#3421) 2025-04-16 10:03:29 +08:00
quick-start-guide.md Update (#2978) 2025-03-23 16:39:35 +08:00
release-notes.md chore: Mass integration of release/0.18 (#3421) 2025-04-16 10:03:29 +08:00
torch.md Update TensorRT-LLM (#2820) 2025-02-25 21:21:49 +08:00