When Operational Load Outgrows the Team

▶ Watch (02:19)

Terraform Enterprise served Booking.com well after the company centralized its IaC effort in 2021. But workspace adoption outpaced the team. Year-over-year growth ran at roughly 4 dozen workspaces annually, reaching 8,000 total by end of 2023. Ticket volume climbed from zero to 300 per year, and the effort to resolve those tickets grew 10x compared to the platform’s early days.

The pattern was linear at best: more workspaces meant more engineers needed to maintain them. At that growth rate, the only alternatives were to keep hiring or to hand infrastructure management to HashiCorp.

The Pre-Migration Architecture and What It Cost

▶ Watch (06:41)

Before the migration, Terraform Enterprise ran on bare-metal servers across two regions, active and standby, with PostgreSQL replication, Redis cache, Prometheus metrics, and S3 state backups. Eighty organizations, one per engineering team, held roughly 100 workspaces each. Maintenance included quarterly compliance reviews, yearly penetration tests, monthly backup drills, quarterly disaster recovery exercises, and manual failover procedures.

Two full-time engineers were already dedicated to keeping this running when the migration decision was made. Every new workspace added surface area to that load without reducing it.

Setting Migration Goals and Controlling Scope

▶ Watch (10:16)

Booking.com set three hard goals for the migration: preserve the existing 80-organization, 8,000-workspace structure for isolation, keep Terraform modules backward-compatible so teams didn’t need to touch their code, and transfer state without any production downtime. The team also made a deliberate decision to freeze scope. Feature requests, including the most-requested one, dynamic organization self-service provisioning, were explicitly deferred until after the migration completed.

That discipline proved important. The migration involved coordinating 80-plus engineering teams across the company. Mixing in improvements would have introduced risk and slipped the deadline.

Execution, Surprises, and Lessons Learned

▶ Watch (16:16)

The migration ran on idempotent scripts: each run either made progress toward full migration or left both workspaces intact on failure, with no cleanup required and no outages for the owning team. Teams migrated on their own schedule after a pilot wave validated the approach. Three surprises surfaced. HCP does not support multiple organizations at Booking.com’s scale, forcing a consolidation to one organization. HCP’s CLI lacked support for managing sensitive variables, adding manual work across all 80 teams. Giving up control of the underlying infrastructure also meant losing observability into platform behavior.

“decisions not for them but with them.” — Balazs Vegvari

Each gap prompted direct engagement with HashiCorp’s product and engineering teams rather than a workaround.

Results: 75% Operational Cost Reduction

▶ Watch (23:59)

Operational costs dropped 75% after the move to HCP. Projected team effort for the year fell 66%, in exchange for two full-time engineers’ development time invested in the migration itself. Booking.com can now double its workspace count without adding operational headcount. The team also freed capacity for work that had been queued for years: automated compliance evidence collection, workspace metadata backup and restore, and native observability into the platform.

“75% in operational cost. This is a win” — Balazs Vegvari

The relationship with HashiCorp shifted during the process. Booking.com’s scale surfaced gaps in the product, and HashiCorp responded by building features that will allow other companies to adopt HCP at similar scale.

Notable Quotes

decisions not for them but with them. Balazs Vegvari · ▶ 21:48

75% in operational cost. This is a win Balazs Vegvari · ▶ 24:11

This was not only a technology challenge Balazs Vegvari · ▶ 14:55

Key Takeaways

  • Ticket volume grew from zero to 300 per year, with resolution effort 10x higher than at launch.
  • Idempotent migration scripts let 80 engineering teams migrate on their own schedule without outages.
  • Moving to HCP cut operational costs 75% and reduced projected team effort 66% in the first year.