When a Mature Platform Still Fails Its Users

▶ Watch (0:27)

Sony’s PlayStation platform team had operators, pipelines, and golden paths. From the outside it looked mature. Yet the backlog kept growing, teams asked for exceptions, others bypassed the platform entirely, and some deployed directly onto specific nodes or clusters. The tension was structural: teams wanted direct infrastructure control while the platform team pushed toward higher abstraction and standardized deployment paths. That gap forced the question from “how do we operate Kubernetes reliably” to “how do we turn Kubernetes into consumable self-service capabilities.”

Yearly Planning Cycles Against Unpredictable Demand

▶ Watch (8:01)

The platform ran on a yearly rhythm: allocate budgets, define projects up front, plan roadmaps around fixed end dates, deliver to milestones. That model worked for years. But every new capability became its own project with a long planning cycle and little room to pivot. Teams finished a project, closed it, and moved on, accumulating systems onto the stack without revisiting whether they were used. The platform became a black box. Requests from users arrived continuously, but the delivery model was designed for annual commitments.

Architecture Patterns That Shaped the Platform Design

▶ Watch (10:42)

Hagen Tonnies drew on two books, “Platform Strategy” by Gregor Hohpe and “Platform Engineering for Architects,” to frame the architecture work. The controller reconciliation loop, observing, analyzing, acting, and repeating, became a guiding pattern. Teams that bypassed it with imperative automation created drift with no clear owner. Tonnies also applied a decompacting principle from Rich Hickey: separating components reveals not just what each does but what makes two components work together. The streaming service achieved error rates below 2% with this SRE practice in place.

Coordination Collapse and the Enablement Team Trap

▶ Watch (20:02)

As more teams joined, one-to-one communication broke down. A single feature release sometimes required five teams in a room. Dependencies surfaced last minute, escalations happened, and simple changes took weeks. The team mapped interaction paths and created an enablement team to help. It backfired: all coordination arrows pointed at that one team. The fix was redrawing capability boundaries using APIs and CRDs, guiding teams to make decisions locally, and aligning them on OKRs rather than requiring constant cross-team meetings. A standing meeting to coordinate work became a signal that boundaries were not well defined.

Measuring the Wrong Things, Then Fixing the Signals

▶ Watch (25:46)

The team could report CPU usage, node health, and API latency for every cluster. They could not answer whether a capability built over six months was adopted or useful. The scoreboard showed green while workarounds multiplied. They replaced yearly commitments with bets framed around user capabilities. New success metrics covered time-to-value, adoption rate, ease of use, and operational efficiency at scale. The definition of done shifted: a feature is done only when it is integrated, documented, supported, and someone actually uses it.

Notable Quotes

and yet somehow your backlog keeps on growing. Um, every team seems to need something slightly different and some ask for exceptions, others bypass your platform entirely. Eugenia Bergman · ▶ 00:49

if a team needs a standing meeting to coordinate work that means that they’re probably their boundaries are not really well defined and that’s not something that we want to have. Eugenia Bergman · ▶ 25:15

done essentially for us now today means that it’s something that someone is able to rely on a given outcome that our users are able to rely on whether that’s unlocking a new capability, whether that’s unlocking or something that is useful for our users. Eugenia Bergman · ▶ 28:11

imperative calls means basically you threw it over the fence like who should now reconcile what you just created as a drift like who who’s doing that Hagen Tonnies · ▶ 13:54

Key Takeaways

  • A platform with operators and golden paths still fails if delivery runs on yearly project cycles.
  • Centralizing coordination through an enablement team solves overload by recreating the bottleneck elsewhere.
  • “Done” means adopted and relied upon, not delivered on scope and closed.

About the Speaker(s)

Eugenia Bergman leads product and delivery practices within Sony Interactive Entertainment’s Platform organisation, helping engineering teams turn infrastructure projects into platform products. Her background spans platform automation, observability, and large-scale delivery, with recent focus on shifting teams from project-oriented to product-oriented ways of working.

Hagen Tonnies is a Staff Platform Architect at Sony Interactive Entertainment, based in Berlin, with 10 years at Sony. His career spans data engineering, software engineering, and platform engineering. He designs hybrid cloud infrastructure for PlayStation’s gaming platforms, specializing in Kubernetes at scale, MLOps, and GPU clusters serving millions of users.