TensorRT-LLMs/docs/source
Robin Kobus cc490de92c
docs: Add KV Cache Management documentation (#3908)
* docs: Add KV Cache Management documentation

* Introduced a new document detailing the hierarchy and event system for KV cache management, including definitions for Pool, Block, and Page.
* Updated the index.rst to include a reference to the new kv-cache-management.md file.

Signed-off-by: Robin Kobus <19427718+Funatiq@users.noreply.github.com>

* Update docs/source/advanced/kv-cache-management.md

Co-authored-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Signed-off-by: Robin Kobus <19427718+Funatiq@users.noreply.github.com>

* Update KV Cache Pool Management

Signed-off-by: Robin Kobus <19427718+Funatiq@users.noreply.github.com>

* docs: Addcross-file links

Signed-off-by: Robin Kobus <19427718+Funatiq@users.noreply.github.com>

* docs: Clarify tokens_per_block

Signed-off-by: Robin Kobus <19427718+Funatiq@users.noreply.github.com>

* docs: Clarify acronyms

Signed-off-by: Robin Kobus <19427718+Funatiq@users.noreply.github.com>

---------

Signed-off-by: Robin Kobus <19427718+Funatiq@users.noreply.github.com>
Co-authored-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
2025-05-21 08:39:28 +02:00
..
_static [Docs] - Reapply #4220 (#4434) 2025-05-19 22:52:37 +08:00
_templates Update TensorRT-LLM (#1725) 2024-06-04 20:26:32 +08:00
advanced docs: Add KV Cache Management documentation (#3908) 2025-05-21 08:39:28 +02:00
architecture chore: Mass Integration 0.19 (#4255) 2025-05-16 10:53:25 +02:00
blogs fix: replace the image links in the blog (#4490) 2025-05-20 22:39:58 +08:00
commands feat: enhance trtllm serve multimodal (#3757) 2025-05-15 16:16:31 -07:00
dev-on-cloud doc: add doc ahout developent on cloud or runpod (#3194) 2025-04-02 18:10:56 +08:00
examples doc: refactor trtllm-serve examples and doc (#3187) 2025-04-04 11:40:43 +08:00
installation [Infra][Docs] - Some clean-up for the CI pipeline and docs (#4419) 2025-05-19 00:07:45 +08:00
llm-api doc: fix path after examples migration (#3814) 2025-04-24 02:36:45 +08:00
media L4 added to readme (#3301) 2025-04-06 19:09:28 +08:00
performance fix: wrong argument name enable_overlap_scheduler (#4433) 2025-05-19 15:02:22 +08:00
python-api Update TensorRT-LLM (#1492) 2024-04-24 14:44:22 +08:00
reference chore: Mass Integration 0.19 (#4255) 2025-05-16 10:53:25 +02:00
torch docs: Add KV Cache Management documentation (#3908) 2025-05-21 08:39:28 +02:00
conf.py [Docs] - Reapply #4220 (#4434) 2025-05-19 22:52:37 +08:00
helper.py doc: refactor trtllm-serve examples and doc (#3187) 2025-04-04 11:40:43 +08:00
index.rst docs: Add KV Cache Management documentation (#3908) 2025-05-21 08:39:28 +02:00
key-features.md Update TensorRT-LLM (#2562) 2024-12-11 00:31:05 -08:00
overview.md chore: Mass integration of release/0.18 (#3421) 2025-04-16 10:03:29 +08:00
quick-start-guide.md chore: Mass Integration 0.19 (#4255) 2025-05-16 10:53:25 +02:00
release-notes.md chore: Mass Integration 0.19 (#4255) 2025-05-16 10:53:25 +02:00
torch.md chore: Mass Integration 0.19 (#4255) 2025-05-16 10:53:25 +02:00