Kubernetes nodes running out of process IDs
What does ZopNight detect here?
ZopNight flags EKS, GKE and AKS nodes whose `PIDPressure` condition is True, meaning `pid.available`, the node's maximum process count minus processes in use, has dropped below the kubelet's eviction threshold. New pods stop landing on the node, and because PIDs have no requests, the kubelet picks pods to evict by priority.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1767 · RC-1867 · RC-1967 |
| Category | reliability |
| Severity | high |
| Metric | PIDPressure node condition |
| Threshold | condition status True |
| Source | ZopNight |
| Permissions used | list nodes · list pods |
Where it applies
Why a node can run out of processes
Every process and thread on Linux consumes a process ID, and the kernel caps how many can exist.
The kubelet tracks the headroom as pid.available, defined on the
node-pressure eviction page
as the maximum PID count minus the current process count. When it falls below an eviction
threshold, the node reports PIDPressure, which means available process identifiers on the Linux
node have run low.
A node without free PIDs cannot start new processes, and the shortage is shared: one pod that
leaks threads can stop every other pod on the machine from forking. The control
plane maps the condition to the node.kubernetes.io/pid-pressure taint, keeping new pods off.
Because PIDs, like inodes, have no requests, the kubelet ranks pods for eviction by priority alone.
Checking nodes for PID exhaustion
kubectl get nodes -o json | jq -r ' .items[] | select(any(.status.conditions[]; .type == "PIDPressure" and .status == "True")) | .metadata.name'kubectl describe node <node-name> shows when the condition last changed, and listing the pods
scheduled on that node narrows down which workload to inspect.
Fires on the condition and nothing else
ZopNight treats the kubelet’s verdict as authoritative. It inspects the node’s collected
conditions and raises the finding only when the PIDPressure entry reads True. It does not
count processes itself or apply a time window, and a node reported with no conditions yields no
finding. When the kubelet clears the condition, the next evaluation clears the recommendation.
Where this finding stops
The pid.available signal is marked Linux-only in the kubelet’s signal table. Memory and disk shortages are covered by
Node under memory pressure
and Node under disk pressure,
and an unresponsive node by
Node not ready.
Risk to every pod on the node
This reliability finding has no savings figure. A PID leak in one container can stop unrelated pods from forking, crash-loop them, and eventually push the node to evict workloads that were behaving correctly.
Containing runaway processes
- Find the pod whose process or thread count keeps climbing, then fix the leak in the application or its restart logic.
- Cap processes per pod with the kubelet’s
--pod-max-pidsflag or thePodPidsLimitsetting in its configuration file, as the process ID limits page describes. - Configure a
pid.availableeviction threshold so the kubelet acts before the node is fully exhausted. - If the node is already unusable, cordon it, drain it and let the node group replace it.