From Broken Party Game to Microservice Architecture on One Pi
The game is an acceleration-based controller battle: PlayStation Move controllers act as spoons balancing eggs, and players knock each other out. Simon brought it to conferences as an icebreaker, but the open source Python game ran on a Raspberry Pi with no monitoring. When something broke, there was no way to know why. His fix was to refactor the game into a microservice architecture with gRPC. That move made OpenTelemetry auto-instrumentation possible and put the whole observability stack, collector, Grafana, and Flagd, onto the same Pi.
Why Prometheus Scraping Missed Entire Games
The game’s only input is motion sensor acceleration. Games with four players can end in 10 seconds. Prometheus scrapes every 60 seconds by default, meaning four to six complete games could finish between polls. Tuning the scrape interval to 10 seconds helped, but going lower made the setup unreliable. Silent poll drops made players feel unresponsive without any visible error. The resulting Grafana graph looked like a square wave with no game signal at all.
Pushing Metrics at 100 Milliseconds
Prometheus also supports push-based metrics, though few teams use it that way. For short-lived game sessions, push fits better than scrape. The team pushed controller manager data every 100 milliseconds and other services less often. That brought reliable end-to-end latency down to 300-500 milliseconds including all overhead. The Grafana graph changed completely. Cautious early play, rising acceleration spikes during combat, and the moment players died all became visible. The same graph at 60-second resolution showed nothing useful.
Cardinality, Label Costs, and Victoria Metrics
The real bottleneck was not cardinality in the traditional sense. Acceleration has three axes and that is all the game needs. The problem was data density. Adding one label, a game ID, to the 32-controller setup multiplied Prometheus series from 1,200 to 18,000. Grafana query time hit 400 milliseconds. With a 500-millisecond dashboard timeout, that left almost no headroom. Victoria Metrics, a drop-in Prometheus replacement, held query times steady under the same load. Every label added to a high-frequency metric needs a cost calculation before it ships.
Feature Flags for Fault Injection and Runtime Control
Testing disconnects in a real venue is impractical. Simon walked from one end of the conference hall to the other to trigger Bluetooth drops, but an empty hall has nearly perfect signal. Instead, the team used OpenFeature flags backed by Flagd to inject faults at runtime. Flags controlled poll drops, full disconnects, and acceleration spikes per controller without restarting the game. Players get color and vibration feedback from controllers, but poll drops are silent. The dashboards, not the game UI, expose those failures.
What CNCF Tooling Still Needs for Real-Time Workloads
OpenTelemetry SDK buffers and OTel Collector batching worked well once configured. The gap is documentation. Every CNCF getting-started guide assumes web services and cloud deployments. IoT and high-frequency data scenarios have no equivalent reference path. The same problem applies to observing large numbers of AI agents at high frequency. The call to action is to run these tools on real-time systems, document what breaks, and contribute that back. The full setup is a Raspberry Pi, a Bluetooth adapter, and PlayStation Move controllers, all reproducible at home.
Notable Quotes
We even tried to run this game with 32 controllers and more on 60 Hz uh on 60 fps. It was no problem at all. Simon Schrottner · ▶ 05:22
out of the box, Prometheus also supports pushing metrics. It’s something that not many people do, but for such a scenario where the data is only collected for a short time, only pushed for a short time, that’s a cool use case. Manuel Timelthaler · ▶ 07:39
adding just one ID to as one additional label to our metrics the game ID suddenly with uh 32 controllers when we uh we would end up instead of 1,200 00 serieses uh in Prometheus with 18,000 serieses Simon Schrottner · ▶ 10:12
port drops, you don’t see. This game has no UI. You have a bit of feedback via the vibration and the color, and you have a bit of feedback via the audio. But if port drops are happening, you don’t know that. But we see that in our dashboards then. Manuel Timelthaler · ▶ 13:06
Key Takeaways
- Prometheus scrape at 60 seconds swallows games shorter than one minute.
- Push-based Prometheus at 100-millisecond intervals achieves 300-500 ms end-to-end latency on a Raspberry Pi.
- One extra label on a 32-controller metric set multiplied Prometheus series 15x, from 1,200 to 18,000.
About the Speaker(s)
Simon Schrottner is a Senior Software Engineer at Dynatrace and an OpenFeature maintainer. He contributes to OpenFeature’s growth in the open source community and drives its adoption internally at Dynatrace.
Manuel Timelthaler is a Software Architect at Tractive, one of the global leaders in health and location tracking for pets. With over a decade of experience across front-end development and cloud architecture, he focuses on building systems that work for both developers and the teams that operate them.