mirror of
https://github.com/vllm-project/vllm.git
synced 2026-08-08 14:58:09 +00:00
Signed-off-by: Florian Woerner <[email protected]> Co-authored-by: Cyrus Leung <[email protected]>
613 B
613 B
llm-d
vLLM can be deployed with llm-d, a Kubernetes-native distributed inference serving stack providing well-lit paths for anyone to serve large generative AI models at scale. It helps achieve the fastest "time to state-of-the-art (SOTA) performance" for key OSS models across most hardware accelerators and infrastructure providers.
You can use vLLM with llm-d directly by following the official guides or via KServe's LLMInferenceService.