TensorRT-LLMs

Fanrong Li 4632a8642d [None][doc] blog: Optimizing DeepSeek-V3.2 on NVIDIA Blackwell GPUs (#10565 ) Signed-off-by: Fanrong Li <23290157+lfr-0531@users.noreply.github.com>	2026-01-09 05:16:00 -05:00
..
Falcon180B-H200_acc.png	Update latest news (#549 )	2023-12-04 22:04:00 +08:00
Falcon180B-H200_DecvOct.png	Update latest news (#549 )	2023-12-04 22:04:00 +08:00
Falcon180B-H200_H200vA100.png	Update latest news (#549 )	2023-12-04 22:04:00 +08:00
Falcon180B-H200_tps.png	Update latest news (#549 )	2023-12-04 22:04:00 +08:00
H200launch_H200vsH100_tps.png	Add Latest News section (#362 )	2023-11-13 15:17:23 +08:00
H200launch_tps.png	Add Latest News section (#365 )	2023-11-13 20:56:22 +08:00
moe_structure.png	Update TensorRT-LLM (#1358 )	2024-03-26 20:47:14 +08:00
tech_blog1_fuse_a_gemm.png	doc： DS r1 min latency blog (#4386 )	2025-05-16 20:20:28 +08:00
tech_blog1_model_details.png	doc： DS r1 min latency blog (#4386 )	2025-05-16 20:20:28 +08:00
tech_blog1_model_overview.png	doc： DS r1 min latency blog (#4386 )	2025-05-16 20:20:28 +08:00
tech_blog1_router_gemm.png	doc： DS r1 min latency blog (#4386 )	2025-05-16 20:20:28 +08:00
tech_blog1_sparse_exp_as_a_gemm.png	doc： DS r1 min latency blog (#4386 )	2025-05-16 20:20:28 +08:00
tech_blog2_acc_relaxed_acceptance.png	draft[doc]: add mtp tech blog (#4580 )	2025-05-23 13:54:21 +08:00
tech_blog2_mtp_eagle.png	[doc] update mtp documents (#5387 )	2025-06-21 16:05:52 +08:00
tech_blog2_mtp_modules.png	draft[doc]: add mtp tech blog (#4580 )	2025-05-23 13:54:21 +08:00
tech_blog2_mtp_vanilla.png	[doc] update mtp documents (#5387 )	2025-06-21 16:05:52 +08:00
tech_blog2_overall_workflow.png	draft[doc]: add mtp tech blog (#4580 )	2025-05-23 13:54:21 +08:00
tech_blog2_perf_and_ar.png	draft[doc]: add mtp tech blog (#4580 )	2025-05-23 13:54:21 +08:00
tech_blog2_relaxed_acceptance.png	draft[doc]: add mtp tech blog (#4580 )	2025-05-23 13:54:21 +08:00
tech_blog2_tree_spec_decoding.png	draft[doc]: add mtp tech blog (#4580 )	2025-05-23 13:54:21 +08:00
tech_blog2_verify_and_accept.png	draft[doc]: add mtp tech blog (#4580 )	2025-05-23 13:54:21 +08:00
tech_blog3_mla_absorb.png	DeepSeek R1 throughut optimization tech blog for Blackwell GPUs (#4791 )	2025-05-30 18:54:19 +08:00
tech_blog4_Picture1.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture2.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture3.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture4.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture5.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture6.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture7.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture8.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture9.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture10.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture11.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture12.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture13.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture14.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture15.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture16.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture17.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture18.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture19.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture20.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture21.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture22.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture23.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture24.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture25.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog5_Picture1.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture2.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture3.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture4.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture5.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture6.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture7.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture8.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture9.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture10.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture11.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture12.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture13.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture14.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture15.png	blog: add qwen3 disagg perf metrics (#5822 )	2025-07-11 16:41:45 +09:00
tech_blog7_accepted_length_case2.png	[doc] Add NGram tech blog (#6311 )	2025-07-25 10:26:33 -07:00
tech_blog7_al_over_iteration_magpie.png	[doc] Add NGram tech blog (#6311 )	2025-07-25 10:26:33 -07:00
tech_blog7_init_sequence_scan.png	[doc] Add NGram tech blog (#6311 )	2025-07-25 10:26:33 -07:00
tech_blog7_magpie_accepted_length_distribution.png	[doc] Add NGram tech blog (#6311 )	2025-07-25 10:26:33 -07:00
tech_blog7_per_token_update.png	[doc] Add NGram tech blog (#6311 )	2025-07-25 10:26:33 -07:00
tech_blog7_speed_up_first_turn.png	[doc] Add NGram tech blog (#6311 )	2025-07-25 10:26:33 -07:00
tech_blog7_speed_up_second_turn.png	[doc] Add NGram tech blog (#6311 )	2025-07-25 10:26:33 -07:00
tech_blog8_communication_kernel.png	[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )	2025-08-01 16:46:15 +08:00
tech_blog8_kernel_breakdown.png	[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )	2025-08-01 16:46:15 +08:00
tech_blog8_moe_aux_kernels1.png	[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )	2025-08-01 16:46:15 +08:00
tech_blog8_moe_aux_kernels2.png	[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )	2025-08-01 16:46:15 +08:00
tech_blog8_perf-1k-1k-dep.png	[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )	2025-08-01 16:46:15 +08:00
tech_blog8_perf-4k-1k-dep.png	[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )	2025-08-01 16:46:15 +08:00
tech_blog8_perf-8k-1k-dep.png	[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )	2025-08-01 16:46:15 +08:00
tech_blog8_perf-8k-1k-e2e-mtp.png	[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )	2025-08-01 16:46:15 +08:00
tech_blog10_baseline_performance_detail.png	[None] [fix] store blog 10 media via lfs (#7375 )	2025-08-30 10:17:53 +08:00
tech_blog10_baseline_performance_overview.png	[None][doc] add adp balance blog (#7213 )	2025-08-28 11:19:34 -04:00
tech_blog10_baseline_round_robin_strategy.png	[None][doc] add adp balance blog (#7213 )	2025-08-28 11:19:34 -04:00
tech_blog10_context_wait_performance.png	[None] [fix] store blog 10 media via lfs (#7375 )	2025-08-30 10:17:53 +08:00
tech_blog10_dataset_token_distribution.png	[None][doc] add adp balance blog (#7213 )	2025-08-28 11:19:34 -04:00
tech_blog10_full_strategy_performance.png	[None] [fix] store blog 10 media via lfs (#7375 )	2025-08-30 10:17:53 +08:00
tech_blog10_tps_ttft_pareto_curve.png	[None][doc] add adp balance blog (#7213 )	2025-08-28 11:19:34 -04:00
tech_blog12_constrained_decoding_pipeline_overlap.png	[None][doc] Tech blog: Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly (#7864 )	2025-09-19 18:38:12 +08:00
tech_blog12_cpu_gpu_synchronization_for_multiple_steps_by_cuda_callback.png	[None][doc] Tech blog: Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly (#7864 )	2025-09-19 18:38:12 +08:00
tech_blog12_cpu_gpu_synchronization_for_multiple_steps.png	[None][doc] Tech blog: Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly (#7864 )	2025-09-19 18:38:12 +08:00
tech_blog12_one_model_vs_two_model.png	[None][doc] Tech blog: Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly (#7864 )	2025-09-19 18:38:12 +08:00
tech_blog12_pareto_curve_json_mode_eval_llama_3.1_8b.png	[None][doc] Tech blog: Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly (#7864 )	2025-09-19 18:38:12 +08:00
tech_blog12_pareto_curve_json_mode_eval_llama_3.3_70b.png	[None][doc] Tech blog: Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly (#7864 )	2025-09-19 18:38:12 +08:00
tech_blog12_pareto_curve_json_schema_bench_llama_3.1_8b.png	[None][doc] Tech blog: Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly (#7864 )	2025-09-19 18:38:12 +08:00
tech_blog12_pareto_curve_json_schema_bench_llama_3.3_70b.png	[None][doc] Tech blog: Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly (#7864 )	2025-09-19 18:38:12 +08:00
tech_blog13_dynasor_demo.gif	[None][doc] scaffolding tech blog part one (#7835 )	2025-09-25 14:41:59 +08:00
tech_blog13_dynasor_hesitation.png	[None][doc] scaffolding tech blog part one (#7835 )	2025-09-25 14:41:59 +08:00
tech_blog13_dynasor_illustration.jpg	[None][doc] scaffolding tech blog part one (#7835 )	2025-09-25 14:41:59 +08:00
tech_blog13_dynasor_pressure_testing.png	[None][doc] scaffolding tech blog part one (#7835 )	2025-09-25 14:41:59 +08:00
tech_blog13_scaffolding_sequence.png	[None][doc] scaffolding tech blog part one (#7835 )	2025-09-25 14:41:59 +08:00
tech_blog14_alltoall_dataflow.png	[None] [blog] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary) (#8323 )	2025-10-13 06:37:17 -07:00
tech_blog14_MTP_parallel_1.png	[None] [blog] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary) (#8323 )	2025-10-13 06:37:17 -07:00
tech_blog14_MTP_parallel_2.png	[None] [blog] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary) (#8323 )	2025-10-13 06:37:17 -07:00
tech_blog14_overview_after_opt.png	[None] [blog] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary) (#8323 )	2025-10-13 06:37:17 -07:00
tech_blog14_overview_before_opt.png	[None] [blog] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary) (#8323 )	2025-10-13 06:37:17 -07:00
tech_blog14_pdloff.png	[None] [blog] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary) (#8323 )	2025-10-13 06:37:17 -07:00
tech_blog14_pdlon.png	[None] [blog] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary) (#8323 )	2025-10-13 06:37:17 -07:00
tech_blog14_perf.png	[None] [blog] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary) (#8323 )	2025-10-13 06:37:17 -07:00
tech_blog15_ds32_wide_ep.png	[None][doc] blog: Optimizing DeepSeek-V3.2 on NVIDIA Blackwell GPUs (#10565 )	2026-01-09 05:16:00 -05:00
tech_blog15_dsa_architecture.png	[None][doc] blog: Optimizing DeepSeek-V3.2 on NVIDIA Blackwell GPUs (#10565 )	2026-01-09 05:16:00 -05:00
tech_blog15_indexer_topk.png	[None][doc] blog: Optimizing DeepSeek-V3.2 on NVIDIA Blackwell GPUs (#10565 )	2026-01-09 05:16:00 -05:00
tech_blog15_radix_select_topk.png	[None][doc] blog: Optimizing DeepSeek-V3.2 on NVIDIA Blackwell GPUs (#10565 )	2026-01-09 05:16:00 -05:00
tp_ep.png	Update TensorRT-LLM (#1358 )	2024-03-26 20:47:14 +08:00
TRT_LLM_v0-5-0_H100vA100_1st.png	Add Latest News section (#315 )	2023-11-08 15:04:33 +08:00
TRT_LLM_v0-5-0_H100vA100_tps.png	Add Latest News section (#315 )	2023-11-08 15:04:33 +08:00
XQA_ThroughputvsLatency.png	Doc update 20240130 (#1009 )	2024-01-31 03:40:22 +08:00

Falcon180B-H200_acc.png

Update latest news (#549 )

2023-12-04 22:04:00 +08:00

Falcon180B-H200_DecvOct.png

Update latest news (#549 )

2023-12-04 22:04:00 +08:00

Falcon180B-H200_H200vA100.png

Update latest news (#549 )

2023-12-04 22:04:00 +08:00

Falcon180B-H200_tps.png

Update latest news (#549 )

2023-12-04 22:04:00 +08:00

H200launch_H200vsH100_tps.png

Add Latest News section (#362 )

2023-11-13 15:17:23 +08:00

H200launch_tps.png

Add Latest News section (#365 )

2023-11-13 20:56:22 +08:00

moe_structure.png

Update TensorRT-LLM (#1358 )

2024-03-26 20:47:14 +08:00

tech_blog1_fuse_a_gemm.png

doc： DS r1 min latency blog (#4386 )

2025-05-16 20:20:28 +08:00

tech_blog1_model_details.png

doc： DS r1 min latency blog (#4386 )

2025-05-16 20:20:28 +08:00

tech_blog1_model_overview.png

doc： DS r1 min latency blog (#4386 )

2025-05-16 20:20:28 +08:00

tech_blog1_router_gemm.png

doc： DS r1 min latency blog (#4386 )

2025-05-16 20:20:28 +08:00

tech_blog1_sparse_exp_as_a_gemm.png

doc： DS r1 min latency blog (#4386 )

2025-05-16 20:20:28 +08:00

tech_blog2_acc_relaxed_acceptance.png

draft[doc]: add mtp tech blog (#4580 )

2025-05-23 13:54:21 +08:00

tech_blog2_mtp_eagle.png

[doc] update mtp documents (#5387 )

2025-06-21 16:05:52 +08:00

tech_blog2_mtp_modules.png

draft[doc]: add mtp tech blog (#4580 )

2025-05-23 13:54:21 +08:00

tech_blog2_mtp_vanilla.png

[doc] update mtp documents (#5387 )

2025-06-21 16:05:52 +08:00

tech_blog2_overall_workflow.png

draft[doc]: add mtp tech blog (#4580 )

2025-05-23 13:54:21 +08:00

tech_blog2_perf_and_ar.png

draft[doc]: add mtp tech blog (#4580 )

2025-05-23 13:54:21 +08:00

tech_blog2_relaxed_acceptance.png

draft[doc]: add mtp tech blog (#4580 )

2025-05-23 13:54:21 +08:00

tech_blog2_tree_spec_decoding.png

draft[doc]: add mtp tech blog (#4580 )

2025-05-23 13:54:21 +08:00

tech_blog2_verify_and_accept.png

draft[doc]: add mtp tech blog (#4580 )

2025-05-23 13:54:21 +08:00

tech_blog3_mla_absorb.png

DeepSeek R1 throughut optimization tech blog for Blackwell GPUs (#4791 )

2025-05-30 18:54:19 +08:00

tech_blog4_Picture1.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture2.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture3.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture4.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture5.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture6.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture7.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture8.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture9.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture10.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture11.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture12.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture13.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture14.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture15.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture16.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture17.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture18.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture19.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture20.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture21.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture22.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture23.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture24.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog4_Picture25.png

blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )

2025-06-05 22:24:04 +08:00

tech_blog5_Picture1.png

blog: Disaggregated Serving in TensorRT-LLM (#5353 )

2025-06-19 18:02:15 +08:00

tech_blog5_Picture2.png

blog: Disaggregated Serving in TensorRT-LLM (#5353 )

2025-06-19 18:02:15 +08:00

tech_blog5_Picture3.png

blog: Disaggregated Serving in TensorRT-LLM (#5353 )

2025-06-19 18:02:15 +08:00

tech_blog5_Picture4.png

blog: Disaggregated Serving in TensorRT-LLM (#5353 )

2025-06-19 18:02:15 +08:00

tech_blog5_Picture5.png

blog: Disaggregated Serving in TensorRT-LLM (#5353 )

2025-06-19 18:02:15 +08:00

tech_blog5_Picture6.png

blog: Disaggregated Serving in TensorRT-LLM (#5353 )

2025-06-19 18:02:15 +08:00

tech_blog5_Picture7.png

blog: Disaggregated Serving in TensorRT-LLM (#5353 )

2025-06-19 18:02:15 +08:00

tech_blog5_Picture8.png

blog: Disaggregated Serving in TensorRT-LLM (#5353 )

2025-06-19 18:02:15 +08:00

tech_blog5_Picture9.png

blog: Disaggregated Serving in TensorRT-LLM (#5353 )

2025-06-19 18:02:15 +08:00

tech_blog5_Picture10.png

blog: Disaggregated Serving in TensorRT-LLM (#5353 )

2025-06-19 18:02:15 +08:00

tech_blog5_Picture11.png

blog: Disaggregated Serving in TensorRT-LLM (#5353 )

2025-06-19 18:02:15 +08:00

tech_blog5_Picture12.png

blog: Disaggregated Serving in TensorRT-LLM (#5353 )

2025-06-19 18:02:15 +08:00

tech_blog5_Picture13.png

blog: Disaggregated Serving in TensorRT-LLM (#5353 )

2025-06-19 18:02:15 +08:00

tech_blog5_Picture14.png

blog: Disaggregated Serving in TensorRT-LLM (#5353 )

2025-06-19 18:02:15 +08:00

tech_blog5_Picture15.png

blog: add qwen3 disagg perf metrics (#5822 )

2025-07-11 16:41:45 +09:00

tech_blog7_accepted_length_case2.png

[doc] Add NGram tech blog (#6311 )

2025-07-25 10:26:33 -07:00

tech_blog7_al_over_iteration_magpie.png

[doc] Add NGram tech blog (#6311 )

2025-07-25 10:26:33 -07:00

tech_blog7_init_sequence_scan.png

[doc] Add NGram tech blog (#6311 )

2025-07-25 10:26:33 -07:00

tech_blog7_magpie_accepted_length_distribution.png

[doc] Add NGram tech blog (#6311 )

2025-07-25 10:26:33 -07:00

tech_blog7_per_token_update.png

[doc] Add NGram tech blog (#6311 )

2025-07-25 10:26:33 -07:00

tech_blog7_speed_up_first_turn.png

[doc] Add NGram tech blog (#6311 )

2025-07-25 10:26:33 -07:00

tech_blog7_speed_up_second_turn.png

[doc] Add NGram tech blog (#6311 )

2025-07-25 10:26:33 -07:00

tech_blog8_communication_kernel.png

[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )

2025-08-01 16:46:15 +08:00

tech_blog8_kernel_breakdown.png

[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )

2025-08-01 16:46:15 +08:00

tech_blog8_moe_aux_kernels1.png

[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )

2025-08-01 16:46:15 +08:00

tech_blog8_moe_aux_kernels2.png

[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )

2025-08-01 16:46:15 +08:00

tech_blog8_perf-1k-1k-dep.png

[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )

2025-08-01 16:46:15 +08:00

tech_blog8_perf-4k-1k-dep.png

[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )

2025-08-01 16:46:15 +08:00

tech_blog8_perf-8k-1k-dep.png

[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )

2025-08-01 16:46:15 +08:00

tech_blog8_perf-8k-1k-e2e-mtp.png

[None][doc] blog: Scaling Expert Parallelism in TensorRT-LLM (Part 2: Performance Status and Optimization) (#6547 )

2025-08-01 16:46:15 +08:00

tech_blog10_baseline_performance_detail.png

[None] [fix] store blog 10 media via lfs (#7375 )

2025-08-30 10:17:53 +08:00

tech_blog10_baseline_performance_overview.png

[None][doc] add adp balance blog (#7213 )

2025-08-28 11:19:34 -04:00

tech_blog10_baseline_round_robin_strategy.png

[None][doc] add adp balance blog (#7213 )

2025-08-28 11:19:34 -04:00

tech_blog10_context_wait_performance.png

[None] [fix] store blog 10 media via lfs (#7375 )

2025-08-30 10:17:53 +08:00

tech_blog10_dataset_token_distribution.png

[None][doc] add adp balance blog (#7213 )

2025-08-28 11:19:34 -04:00

tech_blog10_full_strategy_performance.png

[None] [fix] store blog 10 media via lfs (#7375 )

2025-08-30 10:17:53 +08:00

tech_blog10_tps_ttft_pareto_curve.png

[None][doc] add adp balance blog (#7213 )

2025-08-28 11:19:34 -04:00

tech_blog12_constrained_decoding_pipeline_overlap.png

[None][doc] Tech blog: Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly (#7864 )

2025-09-19 18:38:12 +08:00

tech_blog12_cpu_gpu_synchronization_for_multiple_steps_by_cuda_callback.png

[None][doc] Tech blog: Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly (#7864 )

2025-09-19 18:38:12 +08:00

tech_blog12_cpu_gpu_synchronization_for_multiple_steps.png

[None][doc] Tech blog: Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly (#7864 )

2025-09-19 18:38:12 +08:00

tech_blog12_one_model_vs_two_model.png

[None][doc] Tech blog: Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly (#7864 )

2025-09-19 18:38:12 +08:00

tech_blog12_pareto_curve_json_mode_eval_llama_3.1_8b.png

[None][doc] Tech blog: Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly (#7864 )

2025-09-19 18:38:12 +08:00

tech_blog12_pareto_curve_json_mode_eval_llama_3.3_70b.png

[None][doc] Tech blog: Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly (#7864 )

2025-09-19 18:38:12 +08:00

tech_blog12_pareto_curve_json_schema_bench_llama_3.1_8b.png

[None][doc] Tech blog: Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly (#7864 )

2025-09-19 18:38:12 +08:00

tech_blog12_pareto_curve_json_schema_bench_llama_3.3_70b.png

[None][doc] Tech blog: Combining Guided Decoding and Speculative Decoding: Making CPU and GPU Cooperate Seamlessly (#7864 )

2025-09-19 18:38:12 +08:00

tech_blog13_dynasor_demo.gif

[None][doc] scaffolding tech blog part one (#7835 )

2025-09-25 14:41:59 +08:00

tech_blog13_dynasor_hesitation.png

[None][doc] scaffolding tech blog part one (#7835 )

2025-09-25 14:41:59 +08:00

tech_blog13_dynasor_illustration.jpg

[None][doc] scaffolding tech blog part one (#7835 )

2025-09-25 14:41:59 +08:00

tech_blog13_dynasor_pressure_testing.png

[None][doc] scaffolding tech blog part one (#7835 )

2025-09-25 14:41:59 +08:00

tech_blog13_scaffolding_sequence.png

[None][doc] scaffolding tech blog part one (#7835 )

2025-09-25 14:41:59 +08:00

tech_blog14_alltoall_dataflow.png

[None] [blog] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary) (#8323 )

2025-10-13 06:37:17 -07:00

tech_blog14_MTP_parallel_1.png

[None] [blog] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary) (#8323 )

2025-10-13 06:37:17 -07:00

tech_blog14_MTP_parallel_2.png

[None] [blog] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary) (#8323 )

2025-10-13 06:37:17 -07:00

tech_blog14_overview_after_opt.png

[None] [blog] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary) (#8323 )

2025-10-13 06:37:17 -07:00

tech_blog14_overview_before_opt.png

[None] [blog] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary) (#8323 )

2025-10-13 06:37:17 -07:00

tech_blog14_pdloff.png

[None] [blog] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary) (#8323 )

2025-10-13 06:37:17 -07:00

tech_blog14_pdlon.png

[None] [blog] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary) (#8323 )

2025-10-13 06:37:17 -07:00

tech_blog14_perf.png

[None] [blog] Scaling Expert Parallelism in TensorRT LLM (Part 3: Pushing the Performance Boundary) (#8323 )

2025-10-13 06:37:17 -07:00

tech_blog15_ds32_wide_ep.png

[None][doc] blog: Optimizing DeepSeek-V3.2 on NVIDIA Blackwell GPUs (#10565 )

2026-01-09 05:16:00 -05:00

tech_blog15_dsa_architecture.png

[None][doc] blog: Optimizing DeepSeek-V3.2 on NVIDIA Blackwell GPUs (#10565 )

2026-01-09 05:16:00 -05:00

tech_blog15_indexer_topk.png

[None][doc] blog: Optimizing DeepSeek-V3.2 on NVIDIA Blackwell GPUs (#10565 )

2026-01-09 05:16:00 -05:00

tech_blog15_radix_select_topk.png

[None][doc] blog: Optimizing DeepSeek-V3.2 on NVIDIA Blackwell GPUs (#10565 )

2026-01-09 05:16:00 -05:00

tp_ep.png

Update TensorRT-LLM (#1358 )

2024-03-26 20:47:14 +08:00

TRT_LLM_v0-5-0_H100vA100_1st.png

Add Latest News section (#315 )

2023-11-08 15:04:33 +08:00

TRT_LLM_v0-5-0_H100vA100_tps.png

Add Latest News section (#315 )

2023-11-08 15:04:33 +08:00

XQA_ThroughputvsLatency.png

Doc update 20240130 (#1009 )

2024-01-31 03:40:22 +08:00