The Race Condition Built Into Kubernetes Node Bootstrap
Kubernetes sets a node’s ready condition to true as soon as kubelet is up and sees a basic CNI config on disk. That bar is low. The CNI plugin may not have finished installing routes. GPU device plugins may not have finished warming up. When the scheduler sends a pod to a node in that state, the result is CrashLoopBackOff and unpredictable startup times. Google’s GKE team hit this repeatedly with ghost failures where a node announced readiness before its infrastructure components had actually finished bootstrapping.
How NRC Gates Scheduling With a Taint-and-Condition Contract
NRC runs as an out-of-band controller in the control plane. When a new node joins, NRC immediately applies a startup taint prefixed with readiness.k8s.io. The scheduler sees that taint and holds pods back. Each infrastructure component, CNI, GPU driver, storage driver, publishes a condition on the node’s status subresource. NRC watches those conditions. Once all required conditions are true, it lifts the taint and the node accepts workloads. The single CRD involved is called NodeReadinessRule. It declares which conditions to watch and which nodes to target via a node selector field.
Two Enforcement Modes and Dry-Run Safety
Bootstrap-only mode adds the taint at node initialization, removes it when conditions are met, then stops monitoring entirely. This fits one-time work like pulling heavy images or initializing GPU drivers. Continuous mode keeps watching throughout the node’s life. If a security agent daemonset crashes mid-day, NRC re-applies the taint until the daemonset recovers. Dry-run mode runs the full evaluation logic but writes results only to a status field, letting operators verify what a rule would do before enforcing it on a live cluster.
Performance at 1,000 Nodes: Sub-35ms Latency, Under 50 MB Memory
The team simulated 1,000 nodes joining simultaneously using Kwok, a Kubernetes-without-kubelet SIG project that avoids real compute cost. NRC tainted all 1,000 nodes, then untainted them after conditions were patched in, without missing a single event. Peak taint-addition velocity hit 500 nodes per second. Controller evaluation logic stayed below 35 milliseconds. API server response time also stayed around 35 milliseconds. Memory never exceeded 50 MB under full load. CPU stayed below a quarter of a core. Go routine count stabilized at 56 and held there after work completed, confirming no goroutine leaks.
Ecosystem Integration: Cluster Autoscaler and Constrained Impersonation
Two integration problems surfaced during early production use. Cluster Autoscaler mistook tainted nodes for unschedulable capacity and started scaling incorrectly. A new --startup-taint-prefix flag added to the Cluster Autoscaler binary lets operators register the readiness.k8s.io prefix at configuration time. Autoscaler then skips those nodes rather than reacting to them. The second problem was blast radius from compromised reporter components. Kubernetes 1.36 ships KEP-5284 constrained impersonation in beta, which limits a reporter’s node-status write access to only the node it runs on, rather than all nodes in the cluster.
Release Status and Roadmap
NRC has shipped three minor releases. v0.3.0 adds the security improvements from constrained impersonation and works with existing Kubernetes clusters without any additional dependencies. That version is planned for use in GKE. Upcoming work includes Helm chart support, a Headlamp UI integration, and splitting the NodeReadinessRule CRD into separate rule and evaluation CRDs to reduce etcd pressure. Currently the controller consumes roughly 100 bytes per node against etcd’s 1.5 MB hard limit, which becomes a concern at very high node counts.
Notable Quotes
This project was born out of real production pain we have felt at Google. Ajay Sundar Karuppasamy · ▶ 02:10
cubernetes nodes have a race condition built into the heart of the bootstrapping process Ajay Sundar Karuppasamy · ▶ 02:34
We want to move away from a hope based scheduling to a world where we know things are ready. Ajay Sundar Karuppasamy · ▶ 05:02
we hit a 500 nodes per second uh taint addition. Karthik K N · ▶ 20:54
we have zero panic, zero kills or zero failures. Karthik K N · ▶ 22:38
Key Takeaways
- NRC blocks scheduling with a
readiness.k8s.io-prefixed taint until all declared component conditions are true. - At 1,000 simultaneous nodes, NRC peaked at 500 taint operations per second below 50 MB of memory.
- v0.3.0 is production-ready, dependency-free, and planned for deployment in GKE.
About the Speaker(s)
Priyanka Saggu is a Kubernetes Engineer at SUSE and Technical Lead for SIG ContribEx. She has contributed to SIG Release, SIG Node, SIG API Machinery, and SIG Testing, and served on the Kubernetes Release Team from v1.23 through v1.31.
Karthik K N is a Cloud Engineer at IBM with three years of contributions to Kubernetes and Cluster API projects, primarily focused on SIG Node.
Sreeram Venkitesh is a Senior Software Engineer on DigitalOcean’s managed Kubernetes team. He is active in SIG Release, SIG Contribex Comms, and SIG Node, and served on the Kubernetes release team from v1.29 to v1.35.
Ajay Sundar Karuppasamy is a Software Engineer on Google’s GKE NodeRuntime team and the creator, subproject owner, and lead maintainer of the Node Readiness Controller. He focuses on SIG Node initiatives covering Swap, memory management, and node bootstrapping.