# Node Under Disk Pressure

> Flags EKS, GKE and AKS nodes reporting DiskPressure, where the kubelet is reclaiming disk and may evict pods.

Source: https://zop.dev/integrations/kubernetes/recommendations/node-under-disk-pressure

---

## What a full node filesystem does to its pods

The kubelet watches filesystem eviction signals: `nodefs.available` and `nodefs.inodesFree` for
the node's root filesystem, plus `imagefs` and `containerfs` equivalents when the node keeps
images or container layers on a separate filesystem. The
[node-pressure eviction page](https://kubernetes.io/docs/concepts/scheduling-eviction/node-pressure-eviction/)
lists the default hard thresholds as `nodefs.available<10%`, `imagefs.available<15%`, and 5% free
inodes on Linux. Meeting any of them makes the kubelet report `DiskPressure`, and the control
plane adds the `node.kubernetes.io/disk-pressure` taint so the scheduler sends new pods elsewhere.

Reclaiming comes next. On a node with a single filesystem, the kubelet garbage-collects dead pods
and containers first, then deletes unused images. If space is still short, it evicts running
pods, ranking them by how much disk they use across local volumes, logs and writable layers.
Those evictions ignore PodDisruptionBudgets, and under a hard threshold the pod gets no graceful
shutdown at all.

## Spotting disk pressure across the fleet

```bash
kubectl get nodes -o json | jq -r '
  .items[]
  | select(any(.status.conditions[]; .type == "DiskPressure" and .status == "True"))
  | .metadata.name'
```

Then list what is running on a flagged node, since the heaviest writers are usually pods with
busy logs or large `emptyDir` volumes:

```bash
kubectl get pods -A --field-selector spec.nodeName=<node-name> -o wide
```

## The kubelet's flag is the whole test

ZopNight reads the conditions it collected for each node and fires when a `DiskPressure` entry
has status `True`. `False` and `Unknown` do not count, and a node with no conditions recorded
produces nothing. ZopNight applies no usage figure or time window of its own; the kubelet's
thresholds decide. Note that the kubelet holds a condition for `eviction-pressure-transition-period`,
5 minutes by default, before lowering it, so the finding can outlast the cleanup by a few minutes.

## Other node problems have their own findings

Memory and process-ID shortages are reported by
[Node under memory pressure](https://zop.dev/integrations/kubernetes/recommendations/node-under-memory-pressure)
and [Node under PID pressure](https://zop.dev/integrations/kubernetes/recommendations/node-under-pid-pressure).
A node that has stopped reporting Ready altogether is covered by
[Node not ready](https://zop.dev/integrations/kubernetes/recommendations/node-not-ready).

## Evicted pods rather than dollars

No saving is attached to this high-severity reliability item. The damage is abrupt pod loss on
the affected node, plus a node that accepts no new work while the taint is in place.

## Relieving the disk

1. Read the condition message in `kubectl describe node <node-name>` to see which filesystem or
   inode count tripped the threshold.
2. Find pods writing heavily to logs or `emptyDir`, and give them `ephemeral-storage` requests and
   limits so the scheduler accounts for their disk use.
3. Trim unused images, and enlarge the boot disk in your node group or node pool template so new
   nodes start with more headroom.
4. If the node cannot recover, cordon it, drain it with `kubectl drain --ignore-daemonsets`, and
   replace it.
