Signed-off-by: ibrahimibrahim <[email protected]> Co-authored-by: ibrahimibrahim <[email protected]> Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
3.9 KiB
llm-d
llm-d is a Kubernetes-native distributed inference framework for serving large language models at scale, with vLLM as its primary inference engine. llm-d coordinates a fleet of vLLM instances across a cluster so that performance holds up under real production traffic, achieving the fastest "time to state-of-the-art (SOTA) performance" for key OSS models across most hardware accelerators.
It is a CNCF Sandbox project founded by Red Hat, Google Cloud, IBM Research, CoreWeave, and NVIDIA.
What llm-d adds to vLLM
A single vLLM server is fast, but at scale the picture changes: across many replicas, cache locality breaks under round-robin load balancing, long prompts inflate time-to-first-token, and accelerators sit underused. llm-d adds the cluster-level layer that vLLM does not aim to provide on its own:
- Prefix-aware routing. Instead of round-robin, llm-d reads vLLM's KV-cache events and routes each request to the replica that already holds its prefix, reusing cache instead of recomputing it.
- Distributed KV-cache management. A global index tracks which token blocks live on which replica, and tiered offloading spills cache to CPU memory or local SSD, extending the working set beyond accelerator HBM.
- Prefill/decode disaggregation. Prompt processing and token generation run on separate vLLM workers, with KV-cache moved over the vLLM NIXL connector, lowering TTFT and steadying per-token latency on long prompts.
- Wide expert-parallelism. Serve large Mixture-of-Experts models such as DeepSeek-R1 and GPT-OSS across nodes with combined data and expert parallelism, for more KV-cache capacity and throughput.
- SLO-aware autoscaling and flow control. Scale vLLM pools on real inference signals (queue depth, true demand) rather than raw GPU utilization, with multi-tenant fairness and priority dispatch.
These are composable. Most teams start by adding prefix-aware routing over an existing vLLM pool, then layer in the rest as specific bottlenecks appear.
Performance
Representative benchmarked results across accelerators:
- 3x higher output throughput and 2x faster TTFT from prefix-aware routing vs round-robin (Llama 3.1 70B, AMD MI300X)
- Up to 70% higher tokens/sec from prefill/decode disaggregation (GPT-OSS, NVIDIA B200)
- 13.9x throughput from hierarchical KV offloading at high concurrency vs GPU-only (NVIDIA H100)
See the full list and reproducible benchmarks on Prism.
Get started
- Deploy the Optimized Baseline with the Quickstart. It stands up an intelligent router over a vLLM pool on Kubernetes in a tested configuration.
- Browse the well-lit path guides, each a tested recipe for one of the capabilities above, and add the optimization that fits your workload.
- Read the Introduction and Architecture overview to see how the pieces wrap your vLLM deployment.
You can also deploy vLLM with llm-d via KServe's LLMInferenceService.
Questions and contributions are welcome on GitHub and Slack.