What Is Wrong with Classic Rate (and What Is Not)

▶ Watch (0:39)

rate() is the most-used PromQL function, counting its siblings increase and delta. The classic implementation looks strictly inside the selected range. That property is easy to reason about, but it forces extrapolation to range boundaries, produces non-integer results, and returns nothing when fewer than two samples land inside the range. A separate heuristic tries to guess series start and end, and when it misfires the results are wrong and hard to diagnose. Composability, useful for distributed PromQL engines, is also absent.

Two New Modes: Anchored and Smoothed

▶ Watch (7:57)

Anchored rate takes the last sample inside the range and the last sample before the range and computes their raw difference. No extrapolation. Integers in, integers out. One sample inside the range is enough. Smoothed rate draws a straight line between the last sample before the range and the first sample inside it, then does the same at the far boundary. The intersection points define the increase. It needs zero samples inside the range, always provides full graphing coverage, and handles missing scrapes well. The tradeoff: smoothed underestimates rate for queries issued at the present moment, so a query offset is required.

Simulated Data: How the Three Methods Compare

▶ Watch (11:31)

Rabenstein built a simulation of 20 tasks serving exactly 123.45 QPS through a rolling restart, where each task stops and a replacement starts. The aggregate rate should be a flat line throughout. Smoothed produces visibly straight lines during the crossover. Anchored lags by up to one full scrape interval (15 seconds in the simulation) because the prior sample can sit anywhere in that window. Classic rate sits somewhere between the two but is influenced by at least three interacting factors. The lowest observed aggregate across all methods was 117.5 QPS with no dropped scrapes.

Resilience Under Data Loss

▶ Watch (17:11)

With 10% of scrapes removed randomly, both classic and anchored become noticeably noisy. Smoothed stays nearly unchanged. When summing all 40 tasks (20 old, 20 new), classic and anchored spread widens to around 104 QPS at the low end. Smoothed holds close to its original line. The heuristic inside classic rate that tries to recover lost samples produces a result that is only marginally higher than the others, making the effort hard to justify. Smoothed handles multiple workers and lost scrapes more gracefully than either alternative.

Why This Took So Long: Open Source Conservatism and the Feature Flag Model

▶ Watch (21:32)

The Prometheus team’s stated principle is that “no is temporary, yes is forever.” Adding seven named alternatives and then discovering half were mistakes would leave the project stuck supporting them. Feature flags help: users who activate one have explicitly opted in and cannot claim accidental adoption. Prometheus remote write shipped as a completely experimental feature and became the foundation of a billion-dollar industry before anyone could revise it. That history shaped the team’s caution. A Prometheus 3.0.1 bug fix broke a brittle subquery workaround that many users had built to approximate anchored rate, and that breakage forced the team to ship a proper implementation.

Q&A

Should anyone still use classic rate now that anchored and smoothed exist? Rabenstein said he would personally choose smoothed, but expects many users to prefer anchored because it replaces their previous brittle workarounds. Deprecating classic rate remains possible but depends on community feedback. ▶ 28:05

Is there a feature flag to make smoothed the default? Rabenstein said that if community sentiment confirms smoothed is the right direction, making it the default is worth considering. ▶ 28:49

When will this land in Mimir? Rabenstein said Grafana users were the loudest about anchored rate because Grafana aligns queries to full minutes, which made the old workaround more reliable for them. Given that Grafana Cloud runs on Mimir, he expects it to arrive soon. ▶ 29:13

Notable Quotes

tlddr nothing Björn Rabenstein · ▶ 01:46

the mantra of open source no is temporary. Yes is forever. Björn Rabenstein · ▶ 22:11

users just use experimental features and uh they are not may maybe they are not even aware it’s experimental and they put it into their critical path and then we cannot really take it away. Björn Rabenstein · ▶ 23:01

most infamous example Prometheus remote right literally billiond dollar industry depending on Prometheus remote right when it was completely experimental Björn Rabenstein · ▶ 23:01

booths are great, right? This is what we are here for, not for talks. Björn Rabenstein · ▶ 29:56

Key Takeaways

  • Anchored rate produces integer results using one sample inside the range and one before it.
  • Smoothed rate stays stable with 10% random scrape loss; classic and anchored do not.
  • Both new modes ship in Prometheus 3.7 behind the promql-extended-range-selectors feature flag.

About the Speaker(s)

Björn Rabenstein, known as Beorn, is a Software Engineer at Grafana Labs and a long-time Prometheus developer. Before Grafana Labs, he was a Production Engineer at SoundCloud, a Site Reliability Engineer at Google, and worked in scientific computing. His brain-dump document on rate calculation trade-offs formed the basis for the design doc that Julian Pivoto distilled and implemented.