LeaderWorkerSet
Multi-node inference with LeaderWorkerSet, including autoscaling and topology-aware placement.
Examples focus on two inference runtimes, vLLM and SGLang, across two APIs.
LeaderWorkerSet covers multi-node deployment patterns with autoscaling and topology-aware placement.
DisaggregatedSet covers disaggregated (prefill/decode) deployments with disagg-aware rollouts, autoscaling, and topology placement.
Multi-node inference with LeaderWorkerSet, including autoscaling and topology-aware placement.
Disaggregated (prefill/decode) inference with DisaggregatedSet, including disagg-aware rollouts, autoscaling, and topology placement.
Was this page helpful?
Glad to hear it! Please tell us how we can improve.
Sorry to hear that. Please tell us how we can improve.