TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

History

Kaiyu Xie 5a5427f86e blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 ) Signed-off-by: juney-nvidia <143764042+juney-nvidia@users.noreply.github.com> Co-authored-by: Xianjie <5410381+qiaoxj07@users.noreply.github.com> Co-authored-by: Dongxu Yang <78518666+dongxuy04@users.noreply.github.com> Co-authored-by: Enwei Zhu <21126786+syuoni@users.noreply.github.com> Co-authored-by: Jun Yang <143764042+juney-nvidia@users.noreply.github.com>		2025-06-05 22:24:04 +08:00
..
.gitkeep	Add Latest News section (#315 )	2023-11-08 15:04:33 +08:00
Falcon180B-H200_acc.png	Update latest news (#549 )	2023-12-04 22:04:00 +08:00
Falcon180B-H200_DecvOct.png	Update latest news (#549 )	2023-12-04 22:04:00 +08:00
Falcon180B-H200_H200vA100.png	Update latest news (#549 )	2023-12-04 22:04:00 +08:00
Falcon180B-H200_tps.png	Update latest news (#549 )	2023-12-04 22:04:00 +08:00
H200launch_H200vsH100_tps.png	Add Latest News section (#362 )	2023-11-13 15:17:23 +08:00
H200launch_tps.png	Add Latest News section (#365 )	2023-11-13 20:56:22 +08:00
moe_structure.png	Update TensorRT-LLM (#1358 )	2024-03-26 20:47:14 +08:00
tech_blog1_fuse_a_gemm.png	doc： DS r1 min latency blog (#4386 )	2025-05-16 20:20:28 +08:00
tech_blog1_model_details.png	doc： DS r1 min latency blog (#4386 )	2025-05-16 20:20:28 +08:00
tech_blog1_model_overview.png	doc： DS r1 min latency blog (#4386 )	2025-05-16 20:20:28 +08:00
tech_blog1_router_gemm.png	doc： DS r1 min latency blog (#4386 )	2025-05-16 20:20:28 +08:00
tech_blog1_sparse_exp_as_a_gemm.png	doc： DS r1 min latency blog (#4386 )	2025-05-16 20:20:28 +08:00
tech_blog2_acc_relaxed_acceptance.png	draft[doc]: add mtp tech blog (#4580 )	2025-05-23 13:54:21 +08:00
tech_blog2_mtp_eagle.png	draft[doc]: add mtp tech blog (#4580 )	2025-05-23 13:54:21 +08:00
tech_blog2_mtp_modules.png	draft[doc]: add mtp tech blog (#4580 )	2025-05-23 13:54:21 +08:00
tech_blog2_mtp_vanilla.png	draft[doc]: add mtp tech blog (#4580 )	2025-05-23 13:54:21 +08:00
tech_blog2_overall_workflow.png	draft[doc]: add mtp tech blog (#4580 )	2025-05-23 13:54:21 +08:00
tech_blog2_perf_and_ar.png	draft[doc]: add mtp tech blog (#4580 )	2025-05-23 13:54:21 +08:00
tech_blog2_relaxed_acceptance.png	draft[doc]: add mtp tech blog (#4580 )	2025-05-23 13:54:21 +08:00
tech_blog2_tree_spec_decoding.png	draft[doc]: add mtp tech blog (#4580 )	2025-05-23 13:54:21 +08:00
tech_blog2_verify_and_accept.png	draft[doc]: add mtp tech blog (#4580 )	2025-05-23 13:54:21 +08:00
tech_blog3_mla_absorb.png	DeepSeek R1 throughut optimization tech blog for Blackwell GPUs (#4791 )	2025-05-30 18:54:19 +08:00
tech_blog4_Picture1.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture2.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture3.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture4.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture5.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture6.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture7.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture8.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture9.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture10.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture11.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture12.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture13.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture14.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture15.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture16.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture17.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture18.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture19.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture20.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture21.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture22.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture23.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture24.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture25.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tp_ep.png	Update TensorRT-LLM (#1358 )	2024-03-26 20:47:14 +08:00
TRT_LLM_v0-5-0_H100vA100_1st.png	Add Latest News section (#315 )	2023-11-08 15:04:33 +08:00
TRT_LLM_v0-5-0_H100vA100_tps.png	Add Latest News section (#315 )	2023-11-08 15:04:33 +08:00
XQA_ThroughputvsLatency.png	Doc update 20240130 (#1009 )	2024-01-31 03:40:22 +08:00