Skip to main content
compliance · gcp

GKE clusters with at least one node pool that has auto-repair turned off

resource types
1
rule IDs covered
1
severity
medium

What does ZopNight detect here?

GKE clusters are flagged when ZopNight records `management.autoRepair` as off on their node pools. With auto-repair on, GKE drains and recreates a node that reports NotReady for roughly 10 minutes; with it off, a broken node sits in the pool, its Pods stay stranded, and someone has to notice it before users do.

Signal and threshold

How ZopNight evaluates GKE clusters with at least one node pool that has auto-repair turned off.
Field Value
Rule IDsRC-1226
Categorycompliance
Severitymedium
Metricnone — pure configuration read
Thresholdnode auto-repair recorded as disabled
SourceZopNight
Permissions usedcontainer.clusters.list · container.clusters.get

How a broken node gets fixed without anyone paging in

Node auto-repair is GKE’s health loop for nodes. Google’s auto-repair page lists the triggers: a node reporting NotReady on consecutive checks for about 10 minutes, a node reporting no status at all for about 10 minutes, or a boot disk out of space for about 30 minutes. GKE then drains the node, waits up to an hour for the drain, and recreates it with the same name.

Turn that off and each of those failures becomes a manual incident. Google also calls out a billing trap: if someone deletes an unhealthy node with kubectl while auto-repair is off, the underlying VM can be orphaned from the cluster and keep being billed. Resizing the node pool is the supported way to remove nodes.

Finding pools with repair disabled

Terminal window
gcloud container node-pools list --cluster=CLUSTER_NAME --location=LOCATION \
--format="table(name,management.autoRepair)"

kubectl get nodes shows which nodes are currently NotReady and would have been repaired.

What the check looks at

ZopNight combines the auto-repair settings of a cluster’s node pools into one cluster-level value during inventory. If any pool has auto-repair off, the cluster is flagged, and the finding is reported against the cluster rather than the individual pool.

When it does not fire

A cluster with no observable node pools, and so no recorded value, is left alone rather than being reported as healthy or unhealthy. Autopilot clusters always repair nodes and cannot turn it off, so in practice this is a Standard cluster finding.

Reliability first, with a cost angle

The saving is $0. The practical costs are Pods that cannot schedule on a dead node and, in the orphaned-VM case above, a VM billed for doing nothing.

Switching auto-repair back on

  1. Enable it per pool: gcloud container node-pools update POOL_NAME --cluster=CLUSTER_NAME --location=LOCATION --enable-autorepair.
  2. Check the pool has spare capacity, or the autoscaler can add a node, so a repair does not leave workloads without room while a node is recreated.
  3. Look in Compute Engine for GKE node VMs that no longer belong to any node pool; after confirming nothing uses them, delete them to stop the charge.
  4. Pair this with GKE Node Auto-Upgrade Not Configured so node versions stay current as well.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·