Six Flavors, 160 Clusters, and a Version Problem

▶ Watch (0:33)

Celonis runs a process intelligence platform across AWS, Azure, and GCP in 22 cloud regions. By 2023 the team operated roughly 160 clusters split across six Kubernetes flavors: kops from 2017, Gardener adopted in 2018, and OpenShift from 2021. None were ever decommissioned. The platform team was consumed by market expansion, upgrades slipped, and many clusters fell completely out of support. The goal was three flavors, one managed service per cloud provider, all on the same version.

Building the Target Platform Before Moving Anything

▶ Watch (6:00)

Design work started in 2023. The guiding principle: every cluster must look the same and expose the same Kubernetes capabilities. Celonis adopted EKS first, then AKS and GKE. Terraform modules with Terragrunt provision all infrastructure. Argo ApplicationSets manage cluster add-ons. Since the project started the team has completed seven upgrade cycles, keeping all clusters on a single version. Karpenter was added later for node management and cost reduction. Control planes are fully managed; Celonis no longer runs its own.

The Wormhole: Zero-Downtime Cross-Cluster Traffic

▶ Watch (8:26)

Migrating workloads independently required that applications on the old and new clusters could communicate transparently. Celonis built the wormhole, a set of Envoy proxies with static configuration using Envoy’s dynamic forward proxy feature. An egress proxy on one cluster forwards HTTP traffic to an ingress proxy on the other. Applications call service X without knowing which cluster hosts it. DNS resolves the target dynamically. Existing service mesh tools like Cilium were not usable because the old clusters ran versions too outdated to support them.

Classifying Workloads and the 40,000-PR Problem

▶ Watch (12:06)

Every application fell into one of three buckets. Stateless apps (90%) scaled up on the new cluster, cut over via the wormhole, then scaled down. Singletons (1%) required scale-down first to avoid two simultaneous replicas, introducing brief but negligible downtime. Stateful apps (9%), including the query engine and ML workloads, needed bespoke one-off procedures. Each migration step produced one or more pull requests. With 10 PRs per app, 100 apps per environment, and 40 environments, the math reached 40,000 PRs total. Manual execution was not viable.

Automation: A CLI That Generates PRs in Batches

▶ Watch (17:34)

Celonis built an internal CLI that automated scale-up, cutover, and scale-down steps and produced pull requests as output, keeping a human in the loop for review. The tool ran batches of 10 or 100 apps at a time. Because output was deterministic, runs were repeatable and consistent. Automation also widened who could execute migrations. Early on only a few specialists could run them, which was its own bottleneck. With the CLI, migrations became self-service across teams, removing the dependency on specialists and speeding up the overall pace.

Lessons: Snowflakes, Scope Creep, and Iterating Forward

▶ Watch (22:33)

The biggest drag on velocity was what the team called snowflakiness. Within a single old flavor, environments differed: some used Linkerd, others Istio; some were HIPAA-hardened; some ran on GovCloud. Each variation forced tooling updates and slowed the migration. The team countered by treating this as a consolidation project, not a refactoring one. Application changes were out of scope. Scope creep requests were refused. Runbooks improved after every migration. Service level objectives helped catch regressions. The rule: never cause the same incident twice.

Q&A

Was tearing down the wormhole disruptive? By teardown time all workloads and ingress had already moved to the new cluster, so no traffic remained on the wormhole. ▶ 26:26

Did the wormhole cause a performance penalty? No measurable penalty was observed, though misconfiguration once caused requests to ping-pong between both sides of the wormhole, collapsing it mid-migration and interrupting connectivity. ▶ 26:53

Why was Gardener decommissioned instead of upgraded? The Gardener clusters were on a very old version, and upgrading hop-by-hop across 160 clusters would have taken longer than the full migration approach. ▶ 28:44

How were rollbacks handled? The team preferred fixing forward but every major migration had rollback procedures in the runbook. Rollbacks were also automated to exit broken states as fast as possible. ▶ 31:23

How is configuration drift prevented post-migration? All changes go through infrastructure as code and git. A single template with feature flags controls what is enabled per cluster, so the same base configuration rolls out everywhere. ▶ 34:01

Notable Quotes

70% of all large scale migrations fail. Jannis Relakis · ▶ 01:05

we have 40,000 PRs in total just to migrate all our apps. This is insane, right? Jannis Relakis · ▶ 17:45

snowflakes are kind of the enemy of repeatability because then suddenly your um your proven procedure doesn’t work anymore and you have to get back to the to the drawing board. Jannis Relakis · ▶ 23:43

we maneuvered ourselves in a position where we couldn’t upgrade anymore. Michael Seiwald-McCarty · ▶ 29:27

we always try to fix forward right but our runbooks uh also included roll back procedures Jannis Relakis · ▶ 31:23

Key Takeaways

  • Six Kubernetes flavors across 160 clusters consolidated to one managed service per cloud provider.
  • A custom Envoy-based wormhole proxy enabled zero-downtime cutover without touching application code.
  • A CLI generating PRs in batches made 40,000 theoretical manual operations automated and repeatable.

About the Speaker(s)

Jannis Relakis is a Senior Platform Engineer at Celonis with over 9 years of experience designing and maintaining cloud-based infrastructure for enterprise applications. His work spans DevOps practices, Kubernetes, containerisation, and CI/CD, with a focus on automating manual processes away.

Michael Seiwald-McCarty is a Staff Platform Engineer at Celonis specializing in production-grade Kubernetes. He has worked in the DevOps field for over 11 years and has been hands-on with Kubernetes since 2017 in a range of capacities, with a current focus on operating Kubernetes at fleet scale.