The Scale Problem Behind Autonomous Driving Inference
Wayve builds a single end-to-end AI model for autonomous driving, trained on data from its own fleet plus partner dashcam OEMs and taxis. That fleet collects thousands of hours of driving data every day. Processing it requires Kubernetes clusters spread across multiple Azure regions, each running more than 1,000 GPU nodes. At peak, the platform handles around 100,000 concurrent workloads spanning different teams, GPU hardware types, priorities, and SLAs.
Why Kubernetes Alone Was Not Enough
Kubernetes scheduler handles pod placement well, but it lacks granular controls for fairness in a competitive multi-tenant environment. During high-churn periods, tens of thousands of pending pods degraded scheduler performance. Smaller teams were effectively starved of GPU capacity by larger ones. Wayve needed a layer that could enforce guaranteed allocations per team while still allowing unused capacity to flow to others without manual intervention.
Kueue: Job Queuing Without Replacing the Scheduler
Kueue is a Kubernetes-native job queuing system that complements, rather than replaces, the Kubernetes scheduler. Wayve scale-tested it to 100,000 pods. Integration required zero code changes to existing workloads. Each team receives a guaranteed resource allocation. When a team leaves capacity unused, Kueue distributes that headroom to others fairly. The scheduler only ever sees pods that can actually land on nodes, which keeps it performant even during heavy bursts.
Results: Utilization Up, Wait Times Down
After launch, average wait times fell across all teams. Smaller teams that had previously been starved saw wait times drop by almost 70%. GPU cluster utilization climbed from 85% to 97%. The Kubernetes scheduler stayed highly performant during burst periods because it no longer processed pods that had no available node capacity. The gains arrived without restructuring workloads or rewriting scheduling logic.
From Decision to Full Production in Under One Month
Wayve went from deciding to adopt Kueue to running it at full production scale in less than a month. That timeline included building out all alerting and monitoring. No code changes were needed in existing workload definitions. For a platform handling 100,000 concurrent jobs across clusters of more than 1,000 GPU nodes each, that deployment speed is notable. Muralikrishnan pointed attendees to Kueue contributor talks at KubeCon for deeper technical details.
Notable Quotes
We operate Kubernetes clusters across multiple Azure regions, each running more than 1,000 GPU nodes. Mukund Muralikrishnan · ▶ 0:35
Especially for the smaller teams which used to get starved, the wait times came down by almost 70%. Mukund Muralikrishnan · ▶ 2:20
we went from desiring to implement Kueue and having it running at full production scale in less than a month, including all alerting and monitoring. Mukund Muralikrishnan · ▶ 2:44
Key Takeaways
- Kueue raised Wayve’s GPU cluster utilization from 85% to 97% with no code changes.
- Smaller teams saw wait times fall by 70% after guaranteed allocations replaced open competition.
- A full production rollout including monitoring took under one month at 100,000-pod scale.
About the Speaker(s)
Mukund Muralikrishnan is a Staff Engineer in the AI Platform org at Wayve, where he focuses on compute and storage infrastructure for large-scale AI and ML training and inference workloads. Earlier at Wayve, he helped build the data platform supporting petabyte-scale data operations.