Kubernetes nodes reporting a state other than Ready
What does ZopNight detect here?
ZopNight flags an EKS, GKE or AKS node whose reported state is anything other than Ready. Kubernetes marks such a node with the `node.kubernetes.io/not-ready` or `node.kubernetes.io/unreachable` taint, stops placing new pods on it, and by default evicts its existing pods after 300 seconds, so the capacity is gone until the node recovers.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1768 · RC-1868 · RC-1968 |
| Category | reliability |
| Severity | high |
| Metric | Ready node condition |
| Threshold | node state other than Ready |
| Source | ZopNight |
| Permissions used | list nodes · get nodes |
Where it applies
A node that stops reporting Ready takes its pods with it
Every node carries a Ready condition in its status. The
node status reference gives it three
values: True when the node is healthy and accepting pods, False when it is unhealthy and not
accepting pods, and Unknown when the node controller has not heard from the node within the
node monitor grace period, 50 seconds by default.
Once Ready sits at False or Unknown past that grace period, the control plane adds a
node.kubernetes.io/not-ready or node.kubernetes.io/unreachable taint and the scheduler stops
assigning new pods there. Pods already on the node receive an automatic toleration of 300
seconds for both taints, as the
taints and tolerations page
explains, so they stay bound for five minutes and are then evicted. DaemonSet pods are the
exception: they tolerate both taints with no time limit.
Listing nodes that are not Ready
kubectl get nodes -o json | jq -r ' .items[] | ([.status.conditions[] | select(.type == "Ready")][0]) as $r | select($r.status != "True") | "\(.metadata.name) Ready=\($r.status) reason=\($r.reason) since=\($r.lastTransitionTime)"'For the full picture on one node, including its addresses, capacity and recent events:
kubectl describe node <node-name>A single status read decides it
ZopNight reads the state it last collected for each node in your EKS, GKE and AKS clusters and fires when that state is anything other than Ready, ignoring letter case. No dwell time or lookback window applies: the finding describes the node as of the most recent discovery, and it clears on the first evaluation after the node reports Ready again.
Nodes this check passes over
A node with no recorded state is skipped rather than guessed at. Nodes that are Ready but starved of a resource are a separate signal, covered by Node under disk pressure, Node under memory pressure and Node under PID pressure. The finding reports the state, not the cause, so it cannot tell a crashed kubelet from a severed network path.
Paying for a node that runs nothing new
This high-severity reliability finding carries no savings estimate. The exposure is lost capacity: the remaining nodes must absorb the evicted pods, and if the virtual machine behind the node is still running, your cloud provider is still billing for it.
Recovering or replacing the node
- Read the
Readycondition’s reason and message inkubectl describe node <node-name>.Unknownmeans the heartbeat stopped arriving, which points at the kubelet, the machine or its network path to the control plane. - Check the kubelet and container runtime on the machine, then its disk, memory and network.
- If it will not recover, run
kubectl cordon <node-name>and thenkubectl drain --ignore-daemonsets <node-name>, which evicts pods while respecting their PodDisruptionBudgets. - Remove the node and let your node group or node pool bring up a replacement.