The Hidden Cost Driving a Million-Dollar AWS Bill
Miro’s AWS billing showed one line item growing at a rate nobody had addressed: intra-region data transfer. When Rodrigo and Iris dug in, 60 to 65% of that cost came from nodes running observability components. The network bill for those components regularly matched or exceeded their compute cost. AWS charges 1 cent per gigabyte in each direction, making it effectively 2 cents per gigabyte for any cross-AZ traffic. At Miro’s scale, that was becoming a million-dollar-per-year problem.
Why Observability Generates So Much Cross-AZ Traffic
Pull-based metrics collectors such as Prometheus and VictoriaMetrics discover targets cluster-wide. They cross AZ boundaries on every scrape with no awareness of zone placement. Then the ingestion pipeline compounds the problem: VM agent, VM insert, and VM storage often land in different zones. Each hop adds billable transfer. Miro’s largest environment handles 140 million time series per day and 12 to 19 gigabytes of uncompressed data per scrape cycle, totaling roughly 25 terabytes per day. At 2 cents per gigabyte, the costs accumulate fast.
Miro’s Non-Negotiables Before Changing the Pipeline
Before touching a metrics pipeline serving 13,000-plus targets, the team set hard constraints: no increase in compute resources, no additional architectural spending, and no degradation in availability or metric quality. They also required centralized storage for alerting and querying to remain intact. The migration ran in two phases to protect reliability. Phase one, making VM agents zone-aware, was a low-risk quick win. Phase two, full AZ isolation across ingestion and querying, is still in progress.
The Technical Implementation: Zone-Aware Scraping
Each AZ gets its own VM agent deployment. A relabeling rule at the end of every scrape job drops targets whose zone annotation does not match the agent’s own zone. Kubernetes 1.35 propagates node topology labels to pods automatically via KEP-4742, now enabled by default. Older clusters can use a mutating webhook, such as Kyverno, to copy labels at pod binding time. The VM agent CRD supports environment variables inside relabeling conditions, so the zone value comes from the Kubernetes downward API with no hard-coded strings.
Handling Topology-Blind Targets
Some targets expose no zone metadata at all. Amazon MSK Kafka brokers, for example, present randomized URLs with no AZ information attached. Miro designates one AZ, specifically the 1A agents, to collect these topology-blind targets. A dedicated AZ placement environment variable creates a logical OR in the relabeling condition: match the real zone value or match an empty string. That covers the blind spots without changing the configuration of agents in other zones. A small amount of cross-AZ traffic for that limited target set is accepted as an intentional tradeoff.
Results and the Phase Two Roadmap
Phase one savings landed in the six figures, from metrics collection alone. End users saw no change at all. Phase two will deploy a fully zone-isolated VictoriaMetrics cluster per AZ, with a federated VM select layer on top so clients can query across zones without any special logic on their side. A GitHub repository with kind-cluster topologies covering multiple observability stack flavors is already public for teams who want to test the blueprint themselves.
Q&A
Did you separately measure ingestion versus query path traffic? They tracked overall cost reduction rather than splitting ingestion from query, and found little room to optimize the read path without downsampling or aggregation, which would break transparent global querying. ▶ 29:26
Notable Quotes
60 sometimes even 65% of that cost was coming from nodes running our observability components Rodrigo Fior Kuntzer · ▶ 02:24
every architecture decision has a cost dimension and you can do something about it. Rodrigo Fior Kuntzer · ▶ 03:39
we can say that our savings were in the six figures. So, we were very happy with the optimizations that we did only for the metrics collection part. Iris Dyrmishi · ▶ 28:49
messing with your metrics pipeline is not a good idea unless you’re sure of what you’re doing. Iris Dyrmishi · ▶ 12:07
Key Takeaways
- AWS cross-AZ transfer costs 2 cents per gigabyte and compounds across every pipeline hop.
- Zone-isolated VM agent deployments with relabeling rules eliminate most cross-AZ scrape traffic.
- Kubernetes 1.35 propagates node topology labels to pods by default, removing the need for webhooks.
About the Speakers
Rodrigo Fior Kuntzer is a Staff Site Reliability Engineer at Miro. He specializes in building high-performance, reliable platforms and has led compute optimization efforts including right-sizing workloads and improving allocation efficiency across Miro’s cloud infrastructure.
Iris Dyrmishi is a Senior Observability Engineer at Miro and a CNCF Ambassador. She focuses on architecting scalable observability platforms for cloud-native systems and helps engineering teams improve their telemetry signals to maintain operational reliability.