Skyscanner’s Scale and Original Mesh Design
Skyscanner runs 600 services across 35,000 pods in four production regions, each with six clusters on AWS, fully on spot via Karpenter. Istio connects it all at 60 million requests per minute. In 2019, Project Globe split large clusters into smaller single-AZ units to reduce blast radius and enable phased upgrades. Each AZ gets two clusters. Connecting those clusters required a custom multicluster design built before Istio’s multicluster support was production-ready, using Istio 1.4 for mTLS and DNS abstraction.
When NLBs Cost 10% of the Cloud Bill
Each cluster had one default gateway backed by an AWS NLB. Six clusters per region meant six NLBs per region, 24 across four regions. The approach was easy to understand and operate. But costs scaled with adoption. As more services joined the platform, those load balancers went from negligible to 10% of Skyscanner’s entire cloud bill. The first-generation mesh had solved connectivity. It had not solved cost.
Replacing NLBs with In-Cluster Egress Gateways
The 2024 migration replaced every NLB with an in-cluster egress gateway. Each cluster got one egress gateway per peer cluster in the region, acting as a layer-4 passthrough with no termination or processing. Service entry IPs were repointed from NLB addresses to these gateway services. Developers saw nothing change. Istiod scaled up 10x, from 3 pods to 30 pods per cluster, because multicluster forces it to repeat its work for each connected cluster. The team considered that an acceptable trade for cutting the 10% cloud bill.
Ambient Mesh: Dropping the Sidecar
Sidecars remained after the NLB removal. With 35,000 pods, one proxy per pod consumed significant CPU and memory. Ambient mesh removes that sidecar. A z-tunnel daemonset runs on each node, handling layer-4 mTLS, authorization, and identity. Layer-7 processing moves to an optional waypoint proxy, shared across a namespace or an entire cluster. Upgrading a sidecar required restarting the whole pod. Upgrading the z-tunnel does not touch the application. Skyscanner deployed one waypoint per cluster after cross-namespace waypoints arrived in Istio 1.27.
Telemetry Without Sidecars: An Unorthodox Solution
Sidecars emitted client and server spans. Those spans fed OpenTelemetry, became metrics, and drove dashboards, alerts, and SLOs. Without sidecars, that pipeline breaks. The fix is unorthodox. Client-side telemetry now comes from the waypoint, which takes the calling pod’s identity from z-tunnel and stamps it on the span. Server-side telemetry comes from the default gateway Skyscanner built in 2019, updated to carry the destination pod’s identity. Latency measurements shift, but the error is under one millisecond.
Lessons From Three Generations of Mesh Migration
The third-generation mesh routes all Skyscanner.io traffic through waypoints. Zero code changes from service owners. Failover, outlier detection, header-based routing, and GitOps subset rollouts all carry over. Istiod is smaller now because it no longer scrapes six peer clusters. Ambient is easy on greenfield. On an existing complex setup, migration takes time. Skyscanner spins new clusters with ambient instead of migrating live ones. Upgrades flow through channels: dev, alpha (3 clusters), beta (8 clusters), then all 24. Argo CD handles syncs in order.
Q&A
How are z-tunnel upgrades handled differently from sidecar upgrades? No problems encountered so far; node restarts cover z-tunnel upgrades without requiring application pod restarts. ▶ 20:02
Does istiod scale horizontally or vertically under multicluster? Horizontally: the cluster ran 3 istiod pods before multicluster and 30 after. ▶ 21:22
Has the z-tunnel migration introduced noticeable latency? No notable raw latency impact found; the trade is three L7 hops (sidecar, gateway, sidecar) for two L7 hops plus two L4 z-tunnels. ▶ 24:43
How are Istio upgrades rolled out across 24 production clusters? Via named channels (dev, alpha=3 clusters, beta=8, then all 24), synced in order by Argo CD with traffic stopped on the first cluster and gradually reintroduced. ▶ 26:10
Notable Quotes
60 million requests a minute. John Clark · ▶ 1:03
But there’s always things to be doing. John Clark · ▶ 9:50
custom glue if you need to make it work. John Clark · ▶ 19:26
it’s built from the ground up in in Rust Steven Thwaites · ▶ 25:39
Key Takeaways
- NLBs consumed 10% of Skyscanner’s cloud bill before the 2024 egress gateway migration.
- Multicluster Istio scaled istiod from 3 pods to 30 pods per cluster to handle added load.
- Ambient mesh telemetry requires custom span identity remapping at the waypoint and default gateway.
- Deploy ambient on new clusters rather than migrating existing complex setups in place.
- One waypoint per cluster became viable after Istio 1.27 added cross-namespace waypoint support.
About the Speaker(s)
John Clark is a senior software engineer at Skyscanner, where he is the subject matter expert on the company’s Istio setup. He has driven improvements both internally and upstream in the open source project, enabling significant gains in mesh efficiency while maintaining resiliency.