TensorRT-LLMs/benchmarks
2026-01-14 21:39:55 -05:00
..
cpp [https://nvbugs/5630196] [fix] Prevent flaky failures in C++ test_e2e.py by using local cached datasets for benchmarking (#10638) 2026-01-14 21:39:55 -05:00
README.md chore: Remove deprecated Python runtime benchmark (#4171) 2025-05-14 18:41:05 +08:00

TensorRT-LLM Benchmarks

Overview

There are currently two workflows to benchmark TensorRT-LLM:

  • trtllm-bench
    • trtllm-bench is native to TensorRT-LLM and is a Python benchmarker for reproducing and testing the performance of TensorRT-LLM.
    • NOTE: This benchmarking suite is a current work in progress and is prone to large changes.
  • C++ benchmarks
    • The recommended workflow that uses TensorRT-LLM C++ API and can take advantage of the latest features of TensorRT-LLM.