The Seven-Day Problem: Why Manual Provisioning Breaks at Scale

▶ Watch (0:15)

DigitalOcean’s platform team was processing 200-plus environment tickets per month. Each one required back-and-forth between the developer team, the platform team, and the security team to settle CPU limits, memory quotas, database access, and firewall rules. The actual provisioning was quick once all that context was gathered, but the gathering took days. End-to-end, developers waited seven days for a staging environment. The platform team had become a support queue, not a product team.

Architecture: Backstage to Argo Events to Argo Workflows to vcluster

▶ Watch (5:19)

Backstage serves as the developer portal. A developer fills in a form, which fires a webhook. Argo Events picks up the webhook and triggers an Argo Workflow. The workflow runs three steps: deploy vcluster, set up the namespace, and verify the cluster. The entire chain is decoupled, so the portal has no knowledge of the provisioning engine and the provisioning engine has no knowledge of the infrastructure beneath it. Everything is declared as code, so changes are versioned and auditable.

vcluster: Isolation Without the Cost of a Full Cluster

▶ Watch (13:52)

Giving each developer a dedicated DOKS cluster would have been expensive. Namespace-level isolation shares CRDs cluster-wide, so one misbehaving developer can affect everyone. vcluster sits between those options: it provides Kubernetes control-plane isolation per developer while the underlying worker nodes stay shared. Developers receive a kubeconfig and interact with what looks like a normal cluster. They never see vcluster. When one worker DOKS cluster reaches capacity, automation spins up an additional cluster with a single command.

Kyverno Guardrails: Policy as the Replacement for Manual Security Review

▶ Watch (15:46)

Kyverno operates at three levels: mutate, generate, and validate. DigitalOcean enforces mandatory CPU and memory limits, trusted container registries, blocked privileged containers, network isolation, ownership labels, and deterministic image versions. The TTL feature is the key cost-control mechanism. Kyverno stamps each vcluster namespace with creation time and TTL duration. A Kubernetes job scans those annotations, calculates expiry, and deletes the environment automatically. No human cleans up anything.

Measured Impact: Tickets Down 90%, Provisioning Under Ten Minutes

▶ Watch (19:51)

Provisioning dropped from seven days to under ten minutes. Infrastructure tickets fell 90%. The remaining tickets cover corner cases and guardrail exceptions. Infrastructure cost fell more than 50% on paper at slide time, and Aparna noted the real figure is higher as adoption has grown since those slides were prepared. Guardrails keep developers from over-provisioning, and TTL cleans up idle environments. Platform engineers now spend time building rather than provisioning.

Notable Quotes

Kubernetes is not the problem, but the operating model is flawed Aparna Prabhu · ▶ 03:08

There has been a 90% reduction in infra tickets. Aparna Prabhu · ▶ 20:16

Tickets don’t scale, but platforms do. Bhavani Indukuri · ▶ 34:08

self-service looks very fancy and everyone wants to implement it, but if you don’t have proper guardrails implemented, then the system become you know, becomes chaos. Bhavani Indukuri · ▶ 04:55

Key Takeaways

  • Seven-day manual provisioning collapsed to under ten minutes using five CNCF tools chained together.
  • Kyverno TTL annotations and a cleanup job eliminate idle environment costs without any human action.
  • A 90% ticket reduction freed the platform team to build features instead of answering provisioning requests.

About the Speaker(s)

Bhavani Indukuri is a Senior Platform Engineer at DigitalOcean, CNCF Ambassador, and KubeCon + CloudNativeCon India 2025 Co-chair. She organises Women in CNCF and focuses on practical platform engineering patterns.

Aparna Prabhu is a Senior Engineering Manager at DigitalOcean leading storage and platform engineering teams. Her work centers on building cloud infrastructure that is both scalable and secure, with a focus on green innovation. She was a keynote speaker at KubeCon India 2024.