Why Regional Clusters Could Not Contain Airbnb’s Blast Radius
By 2023, Airbnb ran roughly 1,000 services on a single regional Kubernetes cluster: 5,000 namespaces, 3,500 deployments, and 36,000 pods. A regional cluster spreads one control plane across multiple AZs and works well, but every infra change or bad deployment risks the entire region. A hypothetical cleanup that accidentally deleted service-mesh resources could take down all cluster traffic. With the CD platform depending on the service mesh, the fix could not be deployed, leaving engineers to hot-patch while pods scaled down under zero load, then collapsed under returning traffic.
The Cell Architecture That Made Zonal Failover Possible
Airbnb introduced two abstractions. A “cell” is a fully self-contained unit: its own Kubernetes cluster, its own service mesh, its own AWS account, and its own rate limits. A “cell set” groups cells across a region. Workloads moved from targeting a hardcoded regional cluster context to specifying a cell set field. That single config change let the platform team reassign services to different clusters without touching application code. Dashboards and CLI tooling updated to expose per-cell metrics, logs, and scale operations alongside region-wide views.
Rebuilding the Deployment System Around Cell Sets
The deployment abstraction called OneTouch had assumed one regional cluster since its creation. Hardcoded cluster contexts appeared in Kubernetes manifests and in CD pipeline configs. The team refactored both. A new cell deploy job replaced the standard deploy job, and the config DSL gained a cell set field. Migration itself became three pipeline stages injected into the existing CD flow: scale up the regional deployment, deploy to cell clusters in parallel, then incrementally shift pods to cells while the regional replica count drains to zero. Each service owner saw the change as a PR review followed by a normal deployment.
Executing at Scale: Batches, Slack Threads, and Omega Services
The mass migration started in early 2025 with 19 services already on cells and 30% of compute migrated. Engineers worked in batches: 10-20 services in parallel for critical workloads, 30-40 for less critical ones. A batch job cloned each repo, auto-migrated configs, opened a PR, looked up the on-call owner from the service catalog, and posted a Slack thread. Migrations ran in three phases: non-prod first, then canary environments, then production. A subset of high-demand “omega” services needed specialized node pools, etcd upgrades from 8 GB to 16 GB, and Friday-night deployment windows.
Results: 95% on Cells, First Zonal Failover Under 10 Minutes
The team hit 50% migrated by June. By year end, 3,000 migration PRs had moved all 1,000 services with zero downtime. Today 95% of Airbnb’s compute runs in cells. Regional clusters are being deleted, confirmed by dashboards where metric collection stops at the deletion timestamp. The payoff arrived in September: a full live zonal traffic shift completed in under 10 minutes with no incidents. Work continues on PR-less workload placement controllers, Headlamp-based cell visualization, and AI-assisted YAML reduction across the organization’s service configs.
Q&A
Is Airbnb running in a single AWS region or multiple? The zonal migration is the first step toward multi-region; some services already run in multiple regions and the expansion is ongoing. ▶ 25:30
Notable Quotes
two migrations 3 months equals 1,000 migrations in 125 years. Sunny Beatteay · ▶ 08:10
step two was just a silent prayer that things didn’t break. Sunny Beatteay · ▶ 09:50
we actually were able to complete our first full zonal traffic migration shift in September 2015 in less than 10 minutes without any instance Sunny Beatteay · ▶ 23:31
probably the most important thing for me is that we’re actually now working on workload placement controllers that will allow for PR-less migrations Sunny Beatteay · ▶ 24:43
Key Takeaways
- A zonal cell architecture limits blast radius to one AZ, enabling traffic failover before root cause is known.
- Injecting migration stages into the existing CD pipeline kept developer effort to roughly half a day per service.
- Three-phase migration (non-prod, canary, production) caught compatibility issues before they reached live traffic.
About the Speaker(s)
Sunny Beatteay is a Senior Infrastructure Engineer at Airbnb specializing in Kubernetes and platform engineering. He builds the systems that let thousands of engineers ship code safely, and occasionally discovers new ways to break them. Outside work he enjoys baking and writing.