Histograms Go Stable and OTLP Gets Faster

▶ Watch (3:25)

Native histograms are officially stable as of 3.8. The feature flag is no longer required, but since 3.9 you must configure histogram ingestion at the scrape level. They compress better than classic histograms and eliminate the indexing overhead of separate bucket series. Kubernetes is a named priority for adoption, since monitoring Kubernetes can spike head series count. On the OTLP side, the receive endpoint now writes directly to the TSDB, delivering a 25% performance improvement. Resource attributes can also be promoted to metric labels.

PromQL Gains Outer Joins and Timestamp Access

▶ Watch (6:13)

Several PromQL extensions are now in experimental status. Outer joins work without awkward workarounds: missing values fill from a user-supplied default. Timestamp extraction is available for interesting samples in a range, so you can pull the timestamp of a peak in the last two hours. Duration arithmetic is on the way. Type and unit labels are a user-visible part of a deeper storage change: metric type and unit will become part of metric identity, eventually allowing metric name overloading. Two new keywords, smoothed and anchored, follow vector selectors.

Start Timestamps Fix Counter Reset Detection

▶ Watch (9:43)

Prometheus historically guessed counter reset times by watching for values lower than the previous reading. Short-lived jobs and long scrape intervals exposed the unreliability of that approach. Start timestamp storage, arriving in 3.11, adds a third value to every sample so Prometheus stores when counting began. This fixes reset detection and enables better OTel compatibility. Alongside this, XOR 2 updates the compression format to encode staleness markers more cleanly and improve read and write speeds. It is testable now but not yet production-stable.

Columnar Storage and a Simpler Histogram Format

▶ Watch (15:30)

Parquet-based columnar storage is maturing in Cortex, Thanos, and Mimir, with prototypes already in production at some companies. Composite sample storage replaces classic histograms with a single sample, cutting indexing cost and making histogram data transactional. A PromQL compatibility shim called HCBS classic is being prototyped so existing queries keep working during migration. OpenMetrics 2.0 is approaching its first release candidate: every histogram becomes one compact line, and native histograms get the same treatment.

AI Spam and a New Governance Model

▶ Watch (26:00)

One person submitted almost 200 contributions in a single week. Almost all were closed as bogus. The maintainers used this as a concrete example of AI-generated PR spam adding review overhead. They asked contributors to disclose AI use, demonstrate understanding of the code they touched, and stay involved after the first merge. The project also ratified a new governance document, with the goal of making contributor onboarding simpler and attracting long-term contributors. Alertmanager gained four new maintainers after the topic was raised at PromCon.

Q&A

Is Thanos still the primary way to manage fleets of Prometheus instances? Bartłomiej Płotka (Thanos co-creator) said Cortex, Thanos, and Mimir share similar architectures with different trade-offs, and none is objectively better than the others. ▶ 29:18

What happened with the metrics renaming work from last year? Schema-based automatic conversion is in the prototype stage; tooling for defining metrics in a schematized form must exist before any prototypes merge into Prometheus. ▶ 30:18

What are the “prom-something” identifiers on the slides? They are pull request numbers in the github.com/prometheus/proposals repo, used as shorthand to reference specific proposals across issues and code comments. ▶ 31:36

Notable Quotes

almost 200 contributions in a week. Jan Fajerski · ▶ 26:37

this is not cool. Jan Fajerski · ▶ 27:00

it’s all about trust. It always was, but Jan Fajerski · ▶ 27:25

the writing the code was always easy, right? Jan Fajerski · ▶ 28:26

Key Takeaways

  • Native histograms are stable in Prometheus 3.8+ and require scrape-level configuration since 3.9.
  • Start timestamp storage arrives in 3.11, replacing guessed counter reset detection with stored facts.
  • One contributor’s almost 200 bogus PRs in a week shows AI spam is a real maintainer burden.

About the Speaker(s)

Bartłomiej Płotka is a Senior Software Engineer at Google working on Cloud Observability. Previously a Principal Software Engineer at Red Hat, he co-founded the CNCF Thanos project and wrote “Efficient Go” with O’Reilly.

Jan Fajerski is a Principal Software Engineer at Red Hat. He came to Prometheus from the Software Defined Storage world, first encountering it through the Ceph community, and has worked full-time on the Prometheus project since 2021.