DaemonSets with ready pods missing from two or more nodes, or over 10% of them
What does ZopNight detect here?
ZopNight flags a Kubernetes DaemonSet on EKS, GKE or AKS when `numberReady` is below `desiredNumberScheduled` and the gap is at least 2 nodes or more than 10% of desired. Nodes without a ready copy lose whatever the DaemonSet provides there, typically logging, monitoring, networking or security agents.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1716 · RC-1816 · RC-1916 |
| Category | reliability |
| Severity | high |
| Metric | status.numberReady vs status.desiredNumberScheduled |
| Threshold | gap of 2 or more nodes, or more than 10% of desired |
| Source | ZopNight |
| Permissions used | list daemonsets.apps · list nodes |
Where it applies
A DaemonSet gap is a per-node blind spot
A DaemonSet is meant to run one pod on every eligible node. Its status reports
desiredNumberScheduled, the number of nodes that should run the pod, and numberReady, the
number of those nodes with a ready copy, as defined in the
DaemonSet API reference.
When the second number trails the first, some nodes run without the agent.
What that means depends on the agent. Missing log shippers lose logs from those nodes, missing metrics agents leave gaps in dashboards, and missing security agents leave workloads on those nodes unwatched. The workloads themselves carry on as if nothing were wrong.
Comparing ready and desired per DaemonSet
kubectl get daemonsets -A -o json | jq -r ' .items[] | select(.status.numberReady < .status.desiredNumberScheduled) | "\(.metadata.namespace)/\(.metadata.name) ready=\(.status.numberReady) desired=\(.status.desiredNumberScheduled) updated=\(.status.updatedNumberScheduled)"'To find the nodes missing a pod, list the DaemonSet’s pods with their nodes and compare against
kubectl get nodes:
kubectl get pods -n <namespace> -l <daemonset-selector> -o wideThe size of gap that counts
ZopNight fires only when desired is above zero and ready is below it, and then only when the gap is meaningful: at least two nodes short, or more than 10% of desired. On a 5-node cluster, one missing pod is 20% and fires. On a 20-node cluster, one missing pod is 5% and does not; two missing pods fire at any size. A missing ready count is read as zero ready.
gap = desiredNumberScheduled - numberReadyfire when gap >= 2 OR gap / desiredNumberScheduled > 0.10Rollouts are left out
If updatedNumberScheduled is below desired, the DaemonSet is in the middle of an update and pods
are being replaced node by node, so ZopNight stays quiet until the update finishes. A DaemonSet
with zero desired nodes, for example one whose nodeSelector matches nothing, is also skipped,
because there is no gap to measure.
Reliability risk rather than savings
This high-severity reliability finding has no dollar figure. The exposure is the service each missing agent was supposed to provide on those nodes.
Tracking down the missing pods
- Run
kubectl describe daemonset <name>for events and the scheduled, ready and available counts. - For nodes with a pod that is not ready, read its events and logs; crash loops and image pull errors are common.
- For nodes with no pod, compare node taints with the DaemonSet’s tolerations. Kubernetes adds
some tolerations automatically, such as for
node.kubernetes.io/not-ready, but custom taints need explicit ones. - Check whether the pod’s requests fit the node’s free capacity with
kubectl describe node <node>, and trim the agent’s requests or free space if they do not.