Skip to main content
reliability · kubernetes

Kubernetes nodes with the MemoryPressure condition set

resource types
1
rule IDs covered
3
severity
high

What does ZopNight detect here?

ZopNight flags EKS, GKE and AKS nodes whose `MemoryPressure` condition is True. The kubelet raises it once `memory.available` falls below an eviction threshold, 100Mi by default on Linux, and starts evicting pods whose memory use exceeds their requests, ahead of those running within them, to keep the node itself alive.

Signal and threshold

How ZopNight evaluates Kubernetes nodes with the MemoryPressure condition set.
Field Value
Rule IDsRC-1766 · RC-1866 · RC-1966
Categoryreliability
Severityhigh
MetricMemoryPressure node condition
Thresholdcondition status True
SourceZopNight
Permissions usedlist nodes · list pods

Low node memory ends in evictions or OOM kills

The kubelet computes memory.available as the node’s memory capacity minus the working set of everything running on it, read from cgroupfs rather than tools like free. According to the node-pressure eviction page, the default hard threshold is memory.available<100Mi on Linux nodes and 500Mi on Windows. When it is crossed, the kubelet reports MemoryPressure and the control plane adds the node.kubernetes.io/memory-pressure taint. Pods in the Guaranteed and Burstable QoS classes get a matching toleration automatically, so in practice the taint keeps new BestEffort pods away.

Eviction order matters here. BestEffort and Burstable pods using more memory than they requested go first, by priority and then by how far over they are. Guaranteed pods, and Burstable pods within their requests, go last. If memory runs out faster than the kubelet can react, the Linux OOM killer takes over, steered by oom_score_adj values of -997 for Guaranteed pods and 1000 for BestEffort ones.

Finding memory-starved nodes

Terminal window
kubectl get nodes -o json | jq -r '
.items[]
| select(any(.status.conditions[]; .type == "MemoryPressure" and .status == "True"))
| .metadata.name'

With a Metrics API provider such as metrics-server installed, rank the pods using the most memory:

Terminal window
kubectl top pod -A --sort-by=memory

One condition, read from the latest snapshot

The finding is driven entirely by the node’s own report. ZopNight looks through the conditions collected for the node, and only a MemoryPressure entry with status True triggers it; nodes without condition data are left out. There is no ZopNight-side memory percentage, so nodes that run hot but stay above the kubelet’s threshold are not flagged.

Container-level memory problems look different

A single container exceeding its own memory limit can be OOM-killed on a perfectly healthy node, and that does not set this condition. Workloads that declare no requests are the likeliest eviction victims, which Missing CPU/memory requests tracks separately. Disk and process shortages have their own pages, starting with Node under disk pressure.

Instability, not a saving

There is no savings figure. The risk is pods killed without warning during peaks, and those restarts cascading onto other nodes that may already be tight.

Bringing memory back under control

  1. Identify the heaviest pods on the node with kubectl top pod and compare usage to requests.
  2. Raise memory requests to match real usage so the scheduler stops overpacking the node, and set memory limits on workloads that grow without bound.
  3. Move memory-heavy workloads to a node group or node pool with larger instances.
  4. Reserve memory for the operating system and Kubernetes daemons with the kubelet’s system-reserved and kube-reserved settings, described in Reserve Compute Resources for System Daemons.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·