Skip to main content
reliability · kubernetes

Kubernetes nodes running out of process IDs

resource types
1
rule IDs covered
3
severity
high

What does ZopNight detect here?

ZopNight flags EKS, GKE and AKS nodes whose `PIDPressure` condition is True, meaning `pid.available`, the node's maximum process count minus processes in use, has dropped below the kubelet's eviction threshold. New pods stop landing on the node, and because PIDs have no requests, the kubelet picks pods to evict by priority.

Signal and threshold

How ZopNight evaluates Kubernetes nodes running out of process IDs.
Field Value
Rule IDsRC-1767 · RC-1867 · RC-1967
Categoryreliability
Severityhigh
MetricPIDPressure node condition
Thresholdcondition status True
SourceZopNight
Permissions usedlist nodes · list pods

Why a node can run out of processes

Every process and thread on Linux consumes a process ID, and the kernel caps how many can exist. The kubelet tracks the headroom as pid.available, defined on the node-pressure eviction page as the maximum PID count minus the current process count. When it falls below an eviction threshold, the node reports PIDPressure, which means available process identifiers on the Linux node have run low.

A node without free PIDs cannot start new processes, and the shortage is shared: one pod that leaks threads can stop every other pod on the machine from forking. The control plane maps the condition to the node.kubernetes.io/pid-pressure taint, keeping new pods off. Because PIDs, like inodes, have no requests, the kubelet ranks pods for eviction by priority alone.

Checking nodes for PID exhaustion

Terminal window
kubectl get nodes -o json | jq -r '
.items[]
| select(any(.status.conditions[]; .type == "PIDPressure" and .status == "True"))
| .metadata.name'

kubectl describe node <node-name> shows when the condition last changed, and listing the pods scheduled on that node narrows down which workload to inspect.

Fires on the condition and nothing else

ZopNight treats the kubelet’s verdict as authoritative. It inspects the node’s collected conditions and raises the finding only when the PIDPressure entry reads True. It does not count processes itself or apply a time window, and a node reported with no conditions yields no finding. When the kubelet clears the condition, the next evaluation clears the recommendation.

Where this finding stops

The pid.available signal is marked Linux-only in the kubelet’s signal table. Memory and disk shortages are covered by Node under memory pressure and Node under disk pressure, and an unresponsive node by Node not ready.

Risk to every pod on the node

This reliability finding has no savings figure. A PID leak in one container can stop unrelated pods from forking, crash-loop them, and eventually push the node to evict workloads that were behaving correctly.

Containing runaway processes

  1. Find the pod whose process or thread count keeps climbing, then fix the leak in the application or its restart logic.
  2. Cap processes per pod with the kubelet’s --pod-max-pids flag or the PodPidsLimit setting in its configuration file, as the process ID limits page describes.
  3. Configure a pid.available eviction threshold so the kubelet acts before the node is fully exhausted.
  4. If the node is already unusable, cordon it, drain it and let the node group replace it.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·