TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-22 11:42:41 +08:00

History

Iman Tabrizian af04b6f6aa bug: Fix hang bug when context server doesn't have enough capacity for KV Cache (#3095 ) * Fix hang bug when KV cache is low Signed-off-by: Iman Tabrizian <itabrizian@nvidia.com> * Review comments Signed-off-by: Iman Tabrizian <itabrizian@nvidia.com> * Fix attentiondp typo Signed-off-by: Iman Tabrizian <itabrizian@nvidia.com> * Add CI test for this case Signed-off-by: Iman Tabrizian <itabrizian@nvidia.com> * fix: Fix the insertion order for responder futures Signed-off-by: Iman Tabrizian <10105175+tabrizian@users.noreply.github.com> * fix: Fix disagg CPP Signed-off-by: Iman Tabrizian <10105175+tabrizian@users.noreply.github.com> --------- Signed-off-by: Iman Tabrizian <itabrizian@nvidia.com> Signed-off-by: Iman Tabrizian <10105175+tabrizian@users.noreply.github.com>		2025-04-21 15:16:55 +08:00
..
dev	Update (#2978 )	2025-03-23 16:39:35 +08:00
qa	test: add llama3.2 ptp test case (#3363 )	2025-04-21 15:15:45 +08:00
test-db	bug: Fix hang bug when context server doesn't have enough capacity for KV Cache (#3095 )	2025-04-21 15:16:55 +08:00
waives.txt	Waive L0 tests (#3709 )	2025-04-21 11:24:00 +08:00