Skip to main content
reliability · kubernetes

Kubernetes nodes with the DiskPressure condition set

resource types
1
rule IDs covered
3
severity
high

What does ZopNight detect here?

ZopNight flags any EKS, GKE or AKS node whose `DiskPressure` condition is True. The kubelet sets it when free disk space or inodes cross an eviction threshold, such as the default hard limit of `nodefs.available<10%`, then garbage-collects dead containers and unused images and, if that is not enough, evicts pods from the node.

Signal and threshold

How ZopNight evaluates Kubernetes nodes with the DiskPressure condition set.
Field Value
Rule IDsRC-1765 · RC-1865 · RC-1965
Categoryreliability
Severityhigh
MetricDiskPressure node condition
Thresholdcondition status True
SourceZopNight
Permissions usedlist nodes · list pods

What a full node filesystem does to its pods

The kubelet watches filesystem eviction signals: nodefs.available and nodefs.inodesFree for the node’s root filesystem, plus imagefs and containerfs equivalents when the node keeps images or container layers on a separate filesystem. The node-pressure eviction page lists the default hard thresholds as nodefs.available<10%, imagefs.available<15%, and 5% free inodes on Linux. Meeting any of them makes the kubelet report DiskPressure, and the control plane adds the node.kubernetes.io/disk-pressure taint so the scheduler sends new pods elsewhere.

Reclaiming comes next. On a node with a single filesystem, the kubelet garbage-collects dead pods and containers first, then deletes unused images. If space is still short, it evicts running pods, ranking them by how much disk they use across local volumes, logs and writable layers. Those evictions ignore PodDisruptionBudgets, and under a hard threshold the pod gets no graceful shutdown at all.

Spotting disk pressure across the fleet

Terminal window
kubectl get nodes -o json | jq -r '
.items[]
| select(any(.status.conditions[]; .type == "DiskPressure" and .status == "True"))
| .metadata.name'

Then list what is running on a flagged node, since the heaviest writers are usually pods with busy logs or large emptyDir volumes:

Terminal window
kubectl get pods -A --field-selector spec.nodeName=<node-name> -o wide

The kubelet’s flag is the whole test

ZopNight reads the conditions it collected for each node and fires when a DiskPressure entry has status True. False and Unknown do not count, and a node with no conditions recorded produces nothing. ZopNight applies no usage figure or time window of its own; the kubelet’s thresholds decide. Note that the kubelet holds a condition for eviction-pressure-transition-period, 5 minutes by default, before lowering it, so the finding can outlast the cleanup by a few minutes.

Other node problems have their own findings

Memory and process-ID shortages are reported by Node under memory pressure and Node under PID pressure. A node that has stopped reporting Ready altogether is covered by Node not ready.

Evicted pods rather than dollars

No saving is attached to this high-severity reliability item. The damage is abrupt pod loss on the affected node, plus a node that accepts no new work while the taint is in place.

Relieving the disk

  1. Read the condition message in kubectl describe node <node-name> to see which filesystem or inode count tripped the threshold.
  2. Find pods writing heavily to logs or emptyDir, and give them ephemeral-storage requests and limits so the scheduler accounts for their disk use.
  3. Trim unused images, and enlarge the boot disk in your node group or node pool template so new nodes start with more headroom.
  4. If the node cannot recover, cordon it, drain it with kubectl drain --ignore-daemonsets, and replace it.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·