Skip to main content
reliability · kubernetes

Kubernetes nodes reporting a state other than Ready

resource types
1
rule IDs covered
3
severity
high

What does ZopNight detect here?

ZopNight flags an EKS, GKE or AKS node whose reported state is anything other than Ready. Kubernetes marks such a node with the `node.kubernetes.io/not-ready` or `node.kubernetes.io/unreachable` taint, stops placing new pods on it, and by default evicts its existing pods after 300 seconds, so the capacity is gone until the node recovers.

Signal and threshold

How ZopNight evaluates Kubernetes nodes reporting a state other than Ready.
Field Value
Rule IDsRC-1768 · RC-1868 · RC-1968
Categoryreliability
Severityhigh
MetricReady node condition
Thresholdnode state other than Ready
SourceZopNight
Permissions usedlist nodes · get nodes

A node that stops reporting Ready takes its pods with it

Every node carries a Ready condition in its status. The node status reference gives it three values: True when the node is healthy and accepting pods, False when it is unhealthy and not accepting pods, and Unknown when the node controller has not heard from the node within the node monitor grace period, 50 seconds by default.

Once Ready sits at False or Unknown past that grace period, the control plane adds a node.kubernetes.io/not-ready or node.kubernetes.io/unreachable taint and the scheduler stops assigning new pods there. Pods already on the node receive an automatic toleration of 300 seconds for both taints, as the taints and tolerations page explains, so they stay bound for five minutes and are then evicted. DaemonSet pods are the exception: they tolerate both taints with no time limit.

Listing nodes that are not Ready

Terminal window
kubectl get nodes -o json | jq -r '
.items[]
| ([.status.conditions[] | select(.type == "Ready")][0]) as $r
| select($r.status != "True")
| "\(.metadata.name) Ready=\($r.status) reason=\($r.reason) since=\($r.lastTransitionTime)"'

For the full picture on one node, including its addresses, capacity and recent events:

Terminal window
kubectl describe node <node-name>

A single status read decides it

ZopNight reads the state it last collected for each node in your EKS, GKE and AKS clusters and fires when that state is anything other than Ready, ignoring letter case. No dwell time or lookback window applies: the finding describes the node as of the most recent discovery, and it clears on the first evaluation after the node reports Ready again.

Nodes this check passes over

A node with no recorded state is skipped rather than guessed at. Nodes that are Ready but starved of a resource are a separate signal, covered by Node under disk pressure, Node under memory pressure and Node under PID pressure. The finding reports the state, not the cause, so it cannot tell a crashed kubelet from a severed network path.

Paying for a node that runs nothing new

This high-severity reliability finding carries no savings estimate. The exposure is lost capacity: the remaining nodes must absorb the evicted pods, and if the virtual machine behind the node is still running, your cloud provider is still billing for it.

Recovering or replacing the node

  1. Read the Ready condition’s reason and message in kubectl describe node <node-name>. Unknown means the heartbeat stopped arriving, which points at the kubelet, the machine or its network path to the control plane.
  2. Check the kubelet and container runtime on the machine, then its disk, memory and network.
  3. If it will not recover, run kubectl cordon <node-name> and then kubectl drain --ignore-daemonsets <node-name>, which evicts pods while respecting their PodDisruptionBudgets.
  4. Remove the node and let your node group or node pool bring up a replacement.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·