The Cost of Paying for Readiness Instead of Execution
Roche’s two platforms hit the same financial wall from different directions. The drug discovery platform generates 150 million predictions and retrains 300 models every month. The clinical data platform processes half a petabyte of GXP trial data. Both ran static node groups that kept infrastructure warm around the clock. A backup job ran once per day, kept its dedicated node busy for one hour, then left it idle for 23. That idle time was pure waste. The fix: provision infrastructure only for job execution and scale to zero between runs.
Karpenter Node Pools: Matching Infrastructure to Workload Shape
Karpenter’s NodePool sets the intent for provisioning. Separate pools for critical, heavy, and GPU workloads let each group carry different rules. The expireAfter parameter defaults to 30 days. Roche set it to 12 or 24 hours depending on the use case, ensuring stale pods do not keep nodes running. For fault-tolerant workloads, Karpenter notifies 2 minutes before spot termination. Node weights let teams mix on-demand and spot in one pool: a fixed on-demand count covers baseline stability, spot instances handle variable load. From Karpenter 1.7, reserved capacity is drawn first automatically.
Consolidation Tuning and Pod Startup Observability
The gap between capacity and utilization is waste. Roche’s target is the industry’s 90% utilization ratio. Switching from cluster autoscaler to Karpenter cut pod pending time from under 5 minutes to roughly 1 minute for 20,000 pods in one environment. Drilling into NodeClaim metrics showed where the 5-minute startup went: 2 seconds to launch the instance, 2 minutes to register the node, 30 seconds for CNI and kubelet, then 2.5 minutes for image pull and app init. Optimizing the AMI, upgrading to Karpenter 1.7, and shrinking the container image reduced total startup time by 50%.
Hidden Costs: Metric Cardinality and Node Churn
Enabling all Karpenter metrics raised the observability platform bill by several thousand dollars. Karpenter tracks over 900 instance types, their availability and cost. Prometheus adds pod and node name labels, creating high-cardinality time series. Datadog and Grafana Cloud charge per series. The fix is metric relabeling and aggregation. Node churn is a separate signal: occasional spikes in node creation and deletion events indicate healthy scaling. Frequent spikes point to over-aggressive consolidation settings. Karpenter’s disruption NodeClaim metric breaks down causes (underutilized, empty, drifted) so teams can pinpoint which policy is responsible.
Event-Based Scaling: OpenFold and Clinical Data Ingestion
Roche runs OpenFold, an open-source protein structure model from the OpenFold Consortium. Input is a chemical sequence supplied by the user. The team wired SQS messages to KEDA, which scales the StatefulSet where calculations run. AWS Fast Snapshot Restore attaches a large database volume on demand. That combination produced a 100% speed-up in calculation time and cut cost 10 times. The clinical data ingestion side handles files from a single kilobyte to gigabytes. A dispatcher reads key-value clinical study profiles and routes jobs to either a standard node or a Karpenter-provisioned high-compute node, accepting a 2.5-minute cold-start latency for long-running processing.
Right-Sizing with Goldilocks and Proving Value
Goldilocks, a CNCF project built on VPA with auto mode disabled, recommends CPU and memory requests without changing anything in the cluster. Karpenter uses pod-level requests to calculate node utilization. Incorrect requests cause two problems: nodes run out of resources, or teams pay for capacity nothing uses. Both Roche platforms reached 90% cost reduction. For the OCEAN clinical platform, the key factor was stopping payment for readiness and starting payment for executions. To prove value to stakeholders, the team tracked DORA metrics alongside FinOps metrics, framing infrastructure provisioning speed as a direct contributor to deployment frequency.
Notable Quotes
paying for readiness instead of executions at and keeping infrastructure ready just in case is waste of money Malgorzata Widelicka · ▶ 02:32
the key factor for us for uh cost reduction was to stop paying for readiness and start paying for executions. Lukasz Ogrodowczyk · ▶ 15:28
observability is not for free, but that’s amazing how much insights you can get from Carpenter metrics. Lukasz Ogrodowczyk · ▶ 24:58
introducing this whole um infrastructure and also adjusting our application logic a little bit comes up with 100% of the speed up of the calculation and also reduce the cost 10 times Malgorzata Widelicka · ▶ 20:39
Key Takeaways
- Scale to zero between job executions; Karpenter cut costs 90% on both Roche platforms.
- Pod startup dropped 50% by upgrading to Karpenter 1.7 and shrinking container images.
- Enable Karpenter metrics selectively; full cardinality raised observability bills by thousands of dollars.
About the Speaker(s)
Malgorzata Widelicka is a DevOps Lead at Roche, operating Data and ML platforms on AWS. She specializes in Kubernetes, automation, and FinOps. She holds a PhD in Physics and is a Certified Kubernetes Application Developer (CKAD).
Lukasz Ogrodowczyk is a DevOps Engineer at Roche working in the pharma R&D division, a highly regulated GXP environment. His expertise covers CI/CD, observability, and Infrastructure as Code. KubeCon EU 2026 is his third consecutive KubeCon Europe appearance.