Scale and Motivation: 600 Shards, 36 Million Operations per Second

▶ Watch (3:10)

Braze runs Valkey as the backbone of its customer engagement platform, handling rate limiting, distributed locks, message deduplication, and Sidekiq queues. As of this talk, Braze runs about 600 highly available shards processing over 36 million operations per second with a total memory capacity of 6.5 terabytes. The platform sends billions, now trillions, of personalized messages per day across channels including email, push, SMS, and even physical letters. Any downtime delays those messages directly.

The Networking Problem: Why EC2-to-Kubernetes Replication is Hard

▶ Watch (7:58)

Kubernetes pod cluster IPs are internal to the cluster. An EC2 Redis primary cannot reach them. A straight RDB file copy requires pausing writes, meaning downtime. Configuring the Kubernetes pod to replicate from EC2 works for zero downtime, but once the pod becomes primary, the EC2 instance cannot replicate back through the internal cluster IP. That kills the rollback path. Both requirements, no downtime and no data loss on rollback, needed to be satisfied at the same time.

The NLB Solution: Routing Layer Four Traffic into the Cluster

▶ Watch (11:50)

The fix is an AWS Network Load Balancer sitting outside the cluster. Each pod gets a dedicated node port (31000 for the server, 31001 for Sentinel). The NLB listener on port 6380 routes to the node port across all Kubernetes nodes. Redis announces its availability at the NLB IP and listener port. After Sentinel promotes the Kubernetes pod to primary, the EC2 instance replicates through the NLB, preserving the rollback path. Two unexpected problems surfaced: estimated NLB data transfer fees exceeded $100,000 per month, and split-brain risk between EC2-side and pod-side Sentinel groups. Adding a seventh Sentinel with a quorum of five resolved the split-brain issue.

Migration Execution: 300 Shards, Zero Data Loss

▶ Watch (19:15)

The migration ran February 2023 through May 2024. A three-phase Helm chart migration mode handled provisioning, replication setup, and cutover. A script ran health checks before each shard migration, verifying replica count, Sentinel count, and config parity. The Kubernetes pod in the same availability zone as the EC2 primary was promoted first to minimize cross-AZ data transfer fees. All 300 shards across seven clusters migrated without data loss. Customers noticed nothing.

Valkey in 6 Weeks: Two Lines Changed

▶ Watch (20:38)

After Redis announced its license change at KubeCon Paris, Braze decided to follow the community to Valkey. The Helm chart change was two lines: image name and version. Valkey is fully backwards compatible with Redis configuration, so no config changes were needed. By this point Braze had grown to 350 shards across 10 clusters. All 350 migrated in 6 weeks. The best-case result was a 90% reduction in P95 latency on one workload type in the busiest cluster, on a like-for-like configuration with no tuning.

What Comes Next: A Community Valkey Operator

▶ Watch (25:32)

Two problems remain. First, primaries from the same shard cluster on the same node, causing uneven CPU and bandwidth usage up to the AWS network bandwidth limit. Pod topology spread constraints can fix this without adding capacity. Second, Braze scales in and out roughly 100 times during a 30-minute talk, but Valkey databases stay static. The in-place pod resizer, now stable in recent Kubernetes versions, could add memory to a running container without pod recreation. To address both, Heyburn’s team is building a community-led Valkey operator targeting cluster mode, standalone, replication, Sentinel, and a cells topology.

Notable Quotes

migrating nearly 300 Redis instances to Kubernetes. And we did all of that without any downtime Joe Heyburn · ▶ 00:17

we estimated that we would have over a hundred thousand dollars a month in NLB data transfer fees just on that. Joe Heyburn · ▶ 17:47

we did nearly 300 shards all in one go across all of our seven clusters, and we did all of that without any data loss Joe Heyburn · ▶ 19:23

All right, I lied. It wasn’t a lot. It was literally just these two lines that we had to change for us to migrate. Joe Heyburn · ▶ 21:13

we saw a 90% reduction in P95 latency just from the migration. Joe Heyburn · ▶ 22:25

Key Takeaways

  • An AWS NLB on port 6380 solved the EC2-to-Kubernetes replication routing problem without downtime.
  • Adding a seventh Sentinel with quorum five prevented split-brain during cross-platform migration windows.
  • Valkey’s Redis backwards compatibility reduced a 350-shard migration to a two-line Helm chart change.

About the Speaker(s)

Joe Heyburn is a Staff Engineer at Braze, serving as technical lead across the In-Memory Databases and Observability teams. He manages Valkey on Kubernetes at scale and builds tooling to support Braze’s high-growth infrastructure.