The Cost of Bigtable: Why Kubebuilder Was Not Enough
Spotify’s Bigtable operator was a success for feature teams. The platform team migrated the entire fleet from an internal autoscaler to Google’s when Google shipped a good enough version, and feature teams never noticed. But building it cost two full-time engineers 8 months. The owning team are no longer comfortable making changes to it today. They had to learn Go, Kubebuilder, controller runtime, K status conventions, and Kubernetes-specific concepts far outside their storage speciality.
28% Duplication and the Iceberg Problem
Spotify has 750 million monthly active users. Every listen generates an event stored in GCS and later copied to BigQuery for SQL queries and dashboarding. That copy pattern means 28% of data in BigQuery is duplicated from GCS. Apache Iceberg plus Google BigLake lets BigQuery query GCS data directly through Iceberg tables, eliminating the copy. The remaining problem was setup: tens to hundreds of lines of YAML, cross-referencing properties, plus an out-of-band API call to register tables in Spotify’s internal Iceberg catalog.
K-pop: Operators Without Kubernetes Knowledge
K-pop, Kubernetes Protobuf Operator, lets platform developers build operators by implementing a gRPC service. Types are expressed in Protobuf and translated into a generated CRD. The reconcile function takes the resource as input and returns status. There is no controller runtime, no client-go, and no Kubernetes knowledge required to write or test the service. The K-pop CLI lets developers run reconcile locally against a file and a gRPC endpoint, seeing exactly how status translates into Kubernetes conventions without ever touching a cluster.
Kro Handles Composition, K-pop Handles Imperative APIs
With a K-pop operator wrapping the Iceberg catalog API call, the data platform team used Kro to compose that resource with the required GCP resources from KCC into a single custom resource, SpotifyIcebergTable. The full complexity of the use case sits under one resource. The platform team accomplished this without writing any Kubernetes operator code. Alexander and Tomas estimate that Kro and K-pop together can cover over 90% of platform developer use cases at Spotify.
Adoption Signal and What Comes Next
Since K-pop launched internally, platform teams that previously resisted extending the declarative infra platform are now reaching out to request onboarding. Spotify recently crossed the 1 million managed resources milestone, a target set at KubeCon North America. Remaining work includes observability at scale, API versioning for K-pop, and versioning of resource graph definitions in Kro. RGD isolation in Kro is also unresolved: today all RGDs are gated through a central config repo because Kro lacks per-RGD permission isolation.
Q&A
Can K-pop decompose a resource into child resources the way native Kubernetes APIs do? K-pop targets low-level resource kinds; Kro handles composition of multiple resources, and the two are meant to be used together. ▶ 22:11
Are there plans to open source K-pop? No official plans exist yet, but both speakers said they personally want to do it once internal development matures. ▶ 28:02
How does Kro compare to Helm? Kro runs entirely in-cluster; Helm is primarily client-side with external templating. They solve different problems even when both are used to compose resources into products. ▶ 24:10
How do you test RGDs? The data platform team built a large end-to-end testing framework from scratch. The team sees that as a capability gap worth closing for future adopters. ▶ 33:16
Notable Quotes
it ended up taking two full-time employees 8 months to get this into production. Alexander Buck · ▶ 08:42
the owning team are no longer comfortable making changes to it. Alexander Buck · ▶ 08:51
28% of the data that we have in BigQuery is duplicated from GCS. Tomas Aschan · ▶ 03:57
K-pop and Crow we think are a very powerful combination. We we think it can solve over 90% of the platform developer use cases within Spotify. Alexander Buck · ▶ 18:22
There aren’t any official plans to open source it. Uh but we personally would love to do that. Alexander Buck · ▶ 28:02
Key Takeaways
- Building a Kubebuilder operator took two engineers 8 months and left the team afraid to change it.
- K-pop lets storage specialists write a gRPC service and never touch Kubernetes operator code.
- Kro plus K-pop together cover an estimated 90% of platform developer use cases at Spotify.
About the Speaker(s)
Alexander Buck is a Platform Engineer on the Core Infrastructure team at Spotify, where he builds the internal declarative infra platform that manages over 1 million cloud resources across the organisation.