TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-02-08 12:12:33 +08:00

History

Neta Zmora c4b36d31ff [#10137 ][feat] AutoDeploy FP8 MoE refactor (#10138 ) The trtllm (cutlass) fp8 moe operator performs W3+W1 fusion (concat) during inference and we want to move this fusion to the model optimization time. The Cutlass MoE kernel is used thru a trtllm torch operator. Its implementation uses two FC operations (fc1 and fc2) while the canonical MoE API defines three GEMM operations and their associated weights (W1, W2, W3) so when we switch from the torch.moe op to the trtllm.moe op we also change terminology from w1, w2, w3 to fc1, fc2. Signed-off-by: Neta Zmora <96238833+nzmora-nvidia@users.noreply.github.com>		2025-12-24 18:58:10 +02:00
..
_torch	[#10137 ][feat] AutoDeploy FP8 MoE refactor (#10138 )	2025-12-24 18:58:10 +02:00
api_stability	[TRTLLM-9654][feat] Support DeepSeek-V32 chat template (#9814 )	2025-12-19 17:05:38 +08:00
bindings	[https://nvbugs/5643631 ][fix] Fix hostfunc seg fault (#10028 )	2025-12-20 07:58:43 -05:00
disaggregated	[TRTLLM-8920][feat] decouple disagg service from fastapi (#8714 )	2025-12-05 10:44:16 +08:00
executor	[https://nvbugs/5720482 ][fix] Fix test rpc streaming (#9902 )	2025-12-13 01:14:43 -08:00
llmapi	[TRTLLM-9737][chore] Add rl perf reproduce script and enhance the robustness of Ray tests (#9939 )	2025-12-24 15:27:01 +08:00
others	[None][chore] Add unittest for otlp tracing (#8716 )	2025-12-09 18:34:08 -08:00
scaffolding	[None][feat] Refactor scaffolding streaming feature and fix openai wo… (#8622 )	2025-10-30 16:02:40 +08:00
tools	[TRTC-102][docs] `--extra_llm_api_options`->`--config` in docs/examples/tests (#10005 )	2025-12-19 13:48:43 -05:00
trt	[TRTLLM-8682][chore] Remove auto_parallel module (#8329 )	2025-10-22 20:53:08 -04:00
utils	[TRTLLM-8376][feat] top-p optimization (removes redundant softmax) (#9411 )	2025-11-25 18:46:48 +01:00
conftest.py	[TRTLLM-9737][chore] Add rl perf reproduce script and enhance the robustness of Ray tests (#9939 )	2025-12-24 15:27:01 +08:00
dump_checkpoint_stats.py	Update TensorRT-LLM (#2936 )	2025-03-18 21:25:19 +08:00
gc_utils.py	[nvbug 5273941] fix: broken cyclic reference detect (#5417 )	2025-07-01 20:12:55 +08:00
profile_utils.py	Update TensorRT-LLM (#2936 )	2025-03-18 21:25:19 +08:00
pytest.ini	[TRTLLM-9181][feat] improve disagg-server prometheus metrics; synchronize workers' clocks when workers are dynamic (#9726 )	2025-12-16 05:16:32 -08:00
test_model_runner_cpp.py	Update TensorRT-LLM (#2936 )	2025-03-18 21:25:19 +08:00
test_pip_install.py	[https://nvbugs/5616189 ][fix] Make more cases use local cached models (#8935 )	2025-11-11 03:14:05 -08:00