Mutation Rate, Not Node Count, Is the Real Scalability Ceiling
Synthetic load tests show the API server sustains 5 to 16 pod mutations per second, and that ceiling holds regardless of how many nodes you add. Nodes beyond that threshold sit idle because the pod write path does not shard. Różacki and Rodrigues frame every subsequent argument around this: the right question is not how large the cluster is but what mutation rate it can sustain. A single misbehaving controller writing the same annotation repeatedly consumed most of the write budget in one Anthropic incident, and the fix required no cluster resize at all.
AI Training Workloads Break Assumptions Built for Microservices
Microservices platforms matured around 2,000 to 4,000 node clusters, stretching toward 10,000. AI training workloads jump straight to tens of thousands of nodes. They run at low pod density and rarely reschedule under steady state, but they require all-or-nothing gang scheduling, a primitive Kubernetes never shipped natively. Failure recovery creates a burst of activity that overwhelms a control plane sized for normal churn. Reinforcement learning and CI workloads add unpredictable spikes. Rodrigues notes that heterogeneity across training, inference, RL, and CI all share one control plane, and hiding that complexity from the rest of Anthropic is his core job.
How Anthropic and Google Built at Tens-of-Thousands-of-Nodes Scale
Anthropic runs on GKE mega clusters backed by Spanner and EKS ultra clusters using a partitioned etcd journaling technology from Amazon. The two control planes require separate runbooks. Anthropic replaced the kube-scheduler with a topology-aware, gang-scheduling custom scheduler, built a custom CoreDNS plugin that watches pods directly to avoid endpoint-slice fanout, and developed its own CRD because StatefulSets enforced batch-and-wait pod creation rates that did not fit at scale. Google’s primary deviation from standard Kubernetes is swapping etcd for Spanner. Everything else feeds back to upstream.
Claude Writing Controllers Cut Delivery from a Quarter to Weeks
In 2023, Anthropic’s infrastructure team was small and filed tickets with CSPs to diagnose cluster problems. By late 2023, hiring grew and the team shipped a first read cache and a first custom scheduler. By 2025, Claude writes most of Anthropic’s infrastructure software. The custom CRD went from design to production in weeks, where the same work previously took at least a quarter. Rodrigues argues the bar for building bespoke Kubernetes controllers will keep dropping, and the community needs a crisper pipeline from bespoke solutions to upstream so that innovations like Kueue and JobSets reach everyone.
Four Architectural Bets Kubernetes Must Win
Clusters will keep growing, possibly reaching a quarter million nodes. Multicluster architectures reduce blast radius and spread control-plane load but add management overhead and demand a centralized routing layer. Self-hosted Kubernetes, today a niche, will grow as AI labs build their own data centers. Bespoke customization will accelerate. Różacki argues the community must preserve Kubernetes’s plug-and-play architecture while building better processes to standardize what innovators build at the edge. The recently announced Workload Autoscaler initiative, backed by all major partners, is one example of that standardization path working.
Specific Upstream Gaps the Community Should Close
The most pressing gaps: controller sharding and work-queue partitioning, because today’s leader-election model makes every controller a single-replica bottleneck. API server backoff signaling so reconnecting clients stop flooding list operations during recovery. Priority-and-fairness profiles for high-churn scenarios so node heartbeats are not queued behind bulk list requests. Workload-aware scheduling improvements landing in 1.35 and 1.36, including preemption fixes that stop a single high-priority pod from stranding thousands of GPUs. Rodrigues also wants admission and scheduling fused in JobSets so pending pods disappear and the spec alone encodes queue depth.
Notable Quotes
in our own synthetic uh uh load tests uh we find that uh it’s in the order of five to 16 uh pod mutations per second obviously that varies a lot depending on what’s the storage layer uh behind kind of that that that that that test. Um and that’s the ceiling on CD uh regardless of the node count. Artur Rodrigues · ▶ 01:57
if you have a control plane that’s falling behind uh you’re literally spending millions of dollars uh uh per day on those ID accelerators. Artur Rodrigues · ▶ 07:10
we have claude roing writing most of our infrastructure software so it’s writing kind of controllers that CRD that I just mentioned it went from a design to actually been in production in a matter of weeks as uh something that uh before would have taken at least a quarter Artur Rodrigues · ▶ 14:03
I don’t see this as a race with a clear winner in terms of like, oh, it’s either very large clusters or many small clusters or self-hosted. Uh each one of those is answering a different set of constraints Artur Rodrigues · ▶ 29:19
Key Takeaways
- Pod mutation rate, capped at 5 to 16 per second, limits scale before node count does.
- Anthropic replaced the kube-scheduler, etcd DNS path, and StatefulSet primitives to reach tens of thousands of nodes.
- Claude writing infrastructure controllers cut a quarter-long delivery to weeks.
- Controller sharding, API server backoff, and preemption fixes are the most pressing upstream needs.
- Clusters will grow toward a quarter million nodes; multicluster and self-hosted paths will both be needed.
About the Speaker(s)
Maciek Różacki is a product manager at Google Cloud responsible for Kubernetes and GKE capabilities covering scalability, HPC, batch, and AI training use cases, with eight years working on large-scale Kubernetes operations.
Artur Rodrigues is a Member of Technical Staff at Anthropic, where he builds and scales the infrastructure behind Anthropic’s AI development. Before Anthropic he held infrastructure roles at Lacework and Meta on large-scale clusters and data systems.