Three Drivers That Force a Platform to Scale
Gayathri Thiyagarajan identified three distinct triggers for scaling. First, user adoption, either organic growth that creeps up slowly or a sudden spike from external demand. Second, the platform itself adds capabilities on top of existing primitives. Third, organizational change such as a merger or consolidation across subdomains. Each driver produces different problems, some technical, some sociotechnical. Building for 10x scale from day one is a bad idea precisely because you cannot know which of these three will hit you first.
Amadeus: 800,000 vCPU and the Monolith Problem
Amadeus migrated from a mainframe to its own data center in Erding, then to Azure. Today it runs 190 Azure subscriptions, 200 clusters, and 800,000 vCPU across 600 applications. The first scaling wall was environment creation speed. The second, harder wall was day-two operations. The platform used a monolithic automation framework where every middleware or database feature had to pass through a central validation cycle and a central rollout team. Delivering updates slowed to a crawl. Anna Kozachenko’s lesson: one golden path is the wrong design.
Splitting the Monolith Into Multiple Golden Paths
Amadeus is now separating core infrastructure deployment from middleware and database deployment so each component has an independent lifecycle. Teams can ship features to users without waiting for a central gate. The principle generalizes: know who your users are, then define a set of golden paths matched to those users rather than one path that everyone must follow. Stéphane Di Cesare added that platforms often reach technical maturity before anyone has confirmed they are solving the problems developers actually have. User research matters as much as architecture.
Measuring Success Beyond Uptime
Amadeus tracks KPIs on automation success rates and environment creation time, and runs a continuous feedback loop with users to adapt the platform in real time. Gayathri Thiyagarajan argued that success metrics need to be defined from day zero. Adoption numbers and end-to-end latencies are the visible signals, but the invisible signal is whether the platform still fits its purpose and where friction hides. Stéphane Di Cesare surveys users directly, asking whether the platform keeps their workflows simple, then follows up on the answers to distinguish product gaps from communication gaps.
Capacity for the Unexpected: The 50% Slack Rule
Stéphane Di Cesare pointed to a pattern seen across many companies: platform teams are pushed to maximize feature delivery with no slack left for operations. Google’s SRE books put the target at roughly 50% unplanned time, and Di Cesare called that a realistic ratio. At AWS, Gayathri Thiyagarajan described a version of the same problem with Kafka: once the platform launched, teams expected the platform team to debug their Kafka applications too. The team had to invest in open source contributions and community learning to build that expertise quickly enough to keep up.
AI Agents Will Push Platform Boundaries Further
Autonomous agents are already generating platform load, and every panelist expected that load to grow. Gayathri Thiyagarajan said platforms need to interface with agents and must be auditable and compliant when something goes wrong. Stéphane Di Cesare noted that DevOps spent ten years optimizing delivery speed; AI now does the same at a larger order of magnitude. The risk is technical debt accumulating at the same rate. Anna Kozachenko’s view: offload code generation, testing, and alert triage to AI so platform teams can focus on product thinking and architecture.
Notable Quotes
I think the the what will be important will be uh not not to uh not to to prevent the technical depth from scaling and this is why things like guard rails will be important. It’s it’s a good time for QA people again now. So five years ago everybody thought QA was dead and now they they’re coming back. Stéphane Di Cesare · ▶ 25:27
This is lesson learned. Don’t invest in one golden pass. Anna Kozachenko · ▶ 12:55
your system is only as fast as the slowest component or is only scalable as far as your least scaled part Gayathri Thiyagarajan · ▶ 27:29
Terraform is not the answer to everything. Anna Kozachenko · ▶ 28:34
Key Takeaways
- A monolithic automation framework blocks day-two operations; split it by component lifecycle.
- Define platform success metrics from day zero, including friction signals, not just uptime.
- Reserve roughly 50% of platform team capacity for unplanned operational work.
About the Speakers
Cat Morris is a Staff Product Manager at Syntasso, working at the intersection of developer experience, infrastructure strategy, and product design. Her background in enterprise platform tooling centers on helping teams treat platforms as products.
Stéphane Di Cesare is a Senior Platform Engineer on the Platform Experience team at DKB, the German online bank with about 5,000 employees. He advocates for platform adoption and focuses on bridging engineering and user needs, drawing on earlier roles at VMware and Accenture.
Anna Kozachenko is a Cloud Platform System Architect at Amadeus, responsible for infrastructure automation, application configuration, and FinOps. She oversees a scope that spans CI/CD tooling, observability, and central orchestration across Amadeus’s 600 applications and 200 platform teams.
Gayathri Thiyagarajan is a Software Engineering Manager at AWS with over 13 years of engineering experience. She scaled Kafka from 10 clusters at Expedia Group to 50,000 clusters and 150,000 brokers at AWS, and currently manages the AWS Resource Explorer service, which handles search across approximately 37 million AWS accounts.