TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-02-13 06:23:57 +08:00

History

nv-guomingz 49044733e1 chore: delete useless gitkeep files. (#6400 ) Signed-off-by: nv-guomingz <137257613+nv-guomingz@users.noreply.github.com>		2025-07-28 11:38:30 -04:00
..
Falcon180B-H200_acc.png
Falcon180B-H200_DecvOct.png
Falcon180B-H200_H200vA100.png
Falcon180B-H200_tps.png
H200launch_H200vsH100_tps.png
H200launch_tps.png
moe_structure.png
tech_blog1_fuse_a_gemm.png
tech_blog1_model_details.png
tech_blog1_model_overview.png
tech_blog1_router_gemm.png
tech_blog1_sparse_exp_as_a_gemm.png
tech_blog2_acc_relaxed_acceptance.png
tech_blog2_mtp_eagle.png	[doc] update mtp documents (#5387 )	2025-06-21 16:05:52 +08:00
tech_blog2_mtp_modules.png
tech_blog2_mtp_vanilla.png	[doc] update mtp documents (#5387 )	2025-06-21 16:05:52 +08:00
tech_blog2_overall_workflow.png
tech_blog2_perf_and_ar.png
tech_blog2_relaxed_acceptance.png
tech_blog2_tree_spec_decoding.png
tech_blog2_verify_and_accept.png
tech_blog3_mla_absorb.png	DeepSeek R1 throughut optimization tech blog for Blackwell GPUs (#4791 )	2025-05-30 18:54:19 +08:00
tech_blog4_Picture1.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture2.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture3.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture4.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture5.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture6.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture7.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture8.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture9.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture10.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture11.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture12.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture13.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture14.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture15.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture16.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture17.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture18.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture19.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture20.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture21.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture22.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture23.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture24.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog4_Picture25.png	blog: Scaling Expert Parallelism in TensorRT-LLM (Part 1: Design and Implementation of Large-scale EP) (#4958 )	2025-06-05 22:24:04 +08:00
tech_blog5_Picture1.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture2.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture3.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture4.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture5.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture6.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture7.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture8.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture9.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture10.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture11.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture12.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture13.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture14.png	blog: Disaggregated Serving in TensorRT-LLM (#5353 )	2025-06-19 18:02:15 +08:00
tech_blog5_Picture15.png	blog: add qwen3 disagg perf metrics (#5822 )	2025-07-11 16:41:45 +09:00
tech_blog7_accepted_length_case2.png	[doc] Add NGram tech blog (#6311 )	2025-07-25 10:26:33 -07:00
tech_blog7_al_over_iteration_magpie.png	[doc] Add NGram tech blog (#6311 )	2025-07-25 10:26:33 -07:00
tech_blog7_init_sequence_scan.png	[doc] Add NGram tech blog (#6311 )	2025-07-25 10:26:33 -07:00
tech_blog7_magpie_accepted_length_distribution.png	[doc] Add NGram tech blog (#6311 )	2025-07-25 10:26:33 -07:00
tech_blog7_per_token_update.png	[doc] Add NGram tech blog (#6311 )	2025-07-25 10:26:33 -07:00
tech_blog7_speed_up_first_turn.png	[doc] Add NGram tech blog (#6311 )	2025-07-25 10:26:33 -07:00
tech_blog7_speed_up_second_turn.png	[doc] Add NGram tech blog (#6311 )	2025-07-25 10:26:33 -07:00
tp_ep.png
TRT_LLM_v0-5-0_H100vA100_1st.png
TRT_LLM_v0-5-0_H100vA100_tps.png
XQA_ThroughputvsLatency.png