GKE clusters with at least one node pool that has auto-repair turned off
What does ZopNight detect here?
GKE clusters are flagged when ZopNight records `management.autoRepair` as off on their node pools. With auto-repair on, GKE drains and recreates a node that reports NotReady for roughly 10 minutes; with it off, a broken node sits in the pool, its Pods stay stranded, and someone has to notice it before users do.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1226 |
| Category | compliance |
| Severity | medium |
| Metric | none — pure configuration read |
| Threshold | node auto-repair recorded as disabled |
| Source | ZopNight |
| Permissions used | container.clusters.list · container.clusters.get |
Where it applies
How a broken node gets fixed without anyone paging in
Node auto-repair is GKE’s health loop for nodes. Google’s auto-repair page lists the triggers: a node reporting NotReady on consecutive checks for about 10 minutes, a node reporting no status at all for about 10 minutes, or a boot disk out of space for about 30 minutes. GKE then drains the node, waits up to an hour for the drain, and recreates it with the same name.
Turn that off and each of those failures becomes a manual incident. Google also calls out a
billing trap: if someone deletes an unhealthy node with kubectl while auto-repair is off, the
underlying VM can be orphaned from the cluster and keep being billed. Resizing the node pool is
the supported way to remove nodes.
Finding pools with repair disabled
gcloud container node-pools list --cluster=CLUSTER_NAME --location=LOCATION \ --format="table(name,management.autoRepair)"kubectl get nodes shows which nodes are currently NotReady and would have been repaired.
What the check looks at
ZopNight combines the auto-repair settings of a cluster’s node pools into one cluster-level value during inventory. If any pool has auto-repair off, the cluster is flagged, and the finding is reported against the cluster rather than the individual pool.
When it does not fire
A cluster with no observable node pools, and so no recorded value, is left alone rather than being reported as healthy or unhealthy. Autopilot clusters always repair nodes and cannot turn it off, so in practice this is a Standard cluster finding.
Reliability first, with a cost angle
The saving is $0. The practical costs are Pods that cannot schedule on a dead node and, in the orphaned-VM case above, a VM billed for doing nothing.
Switching auto-repair back on
- Enable it per pool:
gcloud container node-pools update POOL_NAME --cluster=CLUSTER_NAME --location=LOCATION --enable-autorepair. - Check the pool has spare capacity, or the autoscaler can add a node, so a repair does not leave workloads without room while a node is recreated.
- Look in Compute Engine for GKE node VMs that no longer belong to any node pool; after confirming nothing uses them, delete them to stop the charge.
- Pair this with GKE Node Auto-Upgrade Not Configured so node versions stay current as well.