Why GPU Workloads Break Standard Serverless Assumptions
GPU workloads violate three assumptions most serverless runtimes make. First, GPUs are not shareable. A single job typically needs the full device and all its memory. Second, the cost of idle compute is real: on AWS or GCP, a GPU instance runs $1.40 to $2.80 per hour. Third, container images for AI models are huge. An 8-gigabyte model plus PyTorch plus CUDA can reach 15 gigabytes, and pulling that over a network takes minutes, not seconds.
The Real-Time Avatar Case: Where Latency Becomes a Product Problem
Cerebrium runs a customer that generates real-time avatars. The pipeline chains text-to-speech, speech-to-text, an LLM, and video generation. Each model has different scaling needs and different GPU memory requirements. But all of them share one constraint: a response in seconds instead of milliseconds kills the product. Misrouting a request to an overloaded pod does not produce CPU throttling. The request simply fails. That makes routing precision far more costly than in any CPU workload.
What Knative Provides and Where It Falls Short for GPU Work
Knative Serving gives Cerebrium request-based autoscaling, scale-to-zero, revision tracking, and an Envoy-backed request buffer out of the box. Vanilla Kubernetes does not. But Knative was designed for lightweight CPU services, not for workloads that are singly concurrent and take minutes per request. On EKS, the Kubernetes informer that feeds pod health to the activator was running 10-second to minute-long delays. Stale health data in a GPU context does not produce quiet failures. It halts all calls to that pod entirely.
Cerebrium’s Modifications to the Knative Activator
To fix stale discovery, Cerebrium modified qproxy to push pod readiness state directly to the activator, bypassing all Kubernetes informer machinery. Propagation became near-instant. They also rebuilt the activator’s internal state model to serialize updates through a single state manager, similar to actor-model mailboxes in the BEAM. For failure handling, pods that drop TCP connections enter a quarantine state and get a chance to recover. Pods with stale IP addresses from the informer trigger an immediate reroute instead of a hard failure. Routing latency at P90 landed at 25 milliseconds.
Global State in Valkey: Solving Activator Sharding
Stock Knative shards the activator so that each instance owns half the targets. When requests take minutes, one activator can saturate while the other sits idle. Cerebrium moved shared routing state into Valkey, which has its own high availability. Every activator now reads from Valkey and can forward to any target, not just its own shard. Lua scripts inside Valkey handle the load balancing decision before the activator proxies the request. The change lets the platform scale to tens of thousands of pods without unnecessary request queuing.
Knative Project Updates: Endpoint Slices, Listener Sets, and NATS JetStream
Knative Serving raised its minimum Kubernetes version to 1.33. Scaling annotations on revisions are now mutable, so changing min-scale no longer requires a full revision rollout. The project plans to switch from regular Endpoints to Endpoint Slices, removing the hard limit of 1,000 endpoints per service. On the eventing side, a community contributor named Andre added NATS JetStream support, available in nightly releases and targeting stable in about three weeks. The Knative Functions CLI is also adding agent-driven deployment and GitOps YAML generation for GitHub Actions and Tekton.
Q&A
How can teams reduce the image pull time for 15-plus-gigabyte GPU containers? Lazy-pull systems such as eStargz and Nydus pull only the metadata immediately and stream image chunks on demand; in practice only about 30% of a 15-gigabyte image is ever read, and GKE image streaming and EKS’s containerd snapshot tweak implement the same idea with a single toggle. ▶ 29:48
Is the Kubernetes kubelet polling cycle itself a cold-start floor? Yes. The kubelet polls for container readiness, so even a 1-second probe interval can add 1.5 seconds to startup; Knative’s qproxy works around this today with aggressive probing, and SIG Node is working on an event-driven alternative from the container runtime. ▶ 31:13
Notable Quotes
if you’re leaving at idle, you’re literally burning money. Elijah Roussos · ▶ 05:02
on AWS uh specifically uh and I’ll get to this in more detail uh the service discovery that Kubernetes by default uh supplies size is glacial. Elijah Roussos · ▶ 12:55
we were experiencing informer delays of upwards of 10 seconds, sometimes even minutes, which means that health information from pods was just out of date. Elijah Roussos · ▶ 14:25
if you just use regular endpoints and that’s how we do the activator um getting into the request path by adjusting endpoints for the K native services um you’re limited to a thousand endpoints. So if you have more than a thousand you’re kind of screwed. Dave Protasowski · ▶ 24:10
Key Takeaways
- GPU routing errors queue requests rather than throttle them, making precision non-negotiable.
- Replacing Kubernetes informer-based discovery with direct qproxy propagation cut routing latency to 25ms at P90.
- Moving activator state into Valkey gives every activator a global view of all targets, eliminating sharding bottlenecks.
About the Speaker(s)
Dave Protasowski is a Knative Steering Committee member and Serving Working Group Lead. He previously worked on Tanzu at VMware/Broadcom and Cloud Foundry at Pivotal.
Elijah Roussos is the Founding Machine Learning Engineer at Cerebrium, a Y Combinator-backed startup focused on serverless GPU infrastructure. He holds degrees in Electrical and Computer Engineering and leads the platform’s ML infrastructure development.