Vitess: One MySQL Interface Across Thousands of Nodes
Vitess starts where every MySQL deployment starts: a single server, probably an RDS instance. Once that server can no longer scale vertically, Vitess partitions the data across shards. Each shard is a standard MySQL replication group with a primary and one or more replicas. Application servers connect through stateless VTGate nodes and still see a single MySQL endpoint. VTGates use Vschema and Vindexes to route each query to the correct shard without any change to application code. The keyspace concept maps one logical database across hundreds of physical nodes.
Who Already Runs on Vitess
Slack messages, GitHub pull request comments, Square payment terminals, and Uber rides all touch Vitess. These deployments run at a scale where a single 500-terabyte RDS instance makes schema changes, backups, and restores operationally painful. Vitess spreads that data across shards, with a rough guideline of 250 GB per shard. Each of these large clusters is managed by a surprisingly small team because Vitess automates the operational work that would otherwise require constant human intervention.
Zero-Downtime Migrations: Move Tables, Re-Shard, and VDiff
The move tables command imports data from a single RDS instance into Vitess shards without downtime. After cutover, a reverse replication stream starts automatically so traffic can switch back if queries slow down tenfold or start throwing errors. Re-shard then splits existing shards further, splitting one hot shard or going from four to eight shards. VDiff closes the loop: it takes a consistent GTID snapshot on source and target, compares every row, and reports any differences. Teams run migrations for as long as needed before committing.
Isolation, Query Coalescing, and Back Pressure
Vitess maps failure domains to cells, typically availability zones. When US-East-1 went down for several hours recently, Vitess deployments spread across other AZs continued serving traffic. VTGate is stateless, so a topology outage does not stop query serving. At the tablet layer, query coalescing watches fingerprints: when a concert ticket sale sends 10,000 identical read queries, only one hits the database and the result fans out to everyone else. Back pressure throttles background jobs like re-shards and schema changes when production load spikes.
Automated Recovery and Zero-Downtime Upgrades
VT Orc runs as a distributed orchestrator and polls tablet health over pub/sub. When a primary fails, VT Orc promotes a replica within seconds. Incremental backups run every five minutes on critical systems, keeping on-disk data never more than five minutes stale. Shard sizes of 500 GB to 2 TB make restore times fast. Failed nodes go into drain mode for offline forensics rather than in-place rebooting. Rolling upgrades work by updating replicas first, then promoting one to primary, with VTGate buffering writes during the brief switchover.
Notable Quotes
whenever you send a Slack message, you know, when you wake up in the morning, for example, that actually gets stored in Vitess. Matt Lord · ▶ 02:15
Instead of 10,000 database queries, you only have one. Rohit Nayak · ▶ 20:19
when you’re talking of four five six nines it’s like seconds of downtime per day or minutes in a year right? Rohit Nayak · ▶ 16:28
the performance and resilience has come in from very deliberate design decisions right at the beginning of the Vitess project. Rohit Nayak · ▶ 26:41
Key Takeaways
- VTGate stateless routing keeps the MySQL interface intact across any number of shards.
- Move tables and re-shard run without downtime and support automatic traffic reversal on problems.
- VDiff uses GTID snapshots to verify every row after a migration completes.
- Query coalescing at VT tablet reduces 10,000 identical reads to a single database hit.
- VT Orc promotes a new primary within seconds of failure using five-minute incremental backups.
About the Speaker(s)
Matt Lord has spent 25+ years in the database space, working as a producer at MySQL and MongoDB and as a consumer at WeWork and Etsy. He is a Vitess maintainer at PlanetScale focused on running MySQL at scale.
Rohit Nayak is a Vitess maintainer at PlanetScale with over three decades of software engineering experience. He leads work on the VReplication module, which powers data replication and migration workflows inside Vitess.