Skip to main content
reliability · kubernetes

PodDisruptionBudgets that currently allow zero evictions

resource types
1
rule IDs covered
3
severity
high

What does ZopNight detect here?

ZopNight flags a PodDisruptionBudget on EKS, GKE or AKS whose `status.disruptionsAllowed` is 0 while `status.expectedPods` is above zero. Every eviction of those pods is then refused with HTTP 429, so `kubectl drain`, node upgrades and cluster scale-downs stall on any node hosting them until the budget allows at least one.

Signal and threshold

How ZopNight evaluates PodDisruptionBudgets that currently allow zero evictions.
Field Value
Rule IDsRC-1747 · RC-1847 · RC-1947
Categoryreliability
Severityhigh
Metricstatus.disruptionsAllowed
Threshold0 allowed disruptions with at least 1 expected pod
SourceZopNight
Permissions usedlist poddisruptionbudgets.policy · list pods

A budget with no room freezes node maintenance

Draining a node goes through the Eviction API, and the API-initiated eviction page spells out the response when a PodDisruptionBudget has no room left: 429 Too Many Requests, the eviction is not currently allowed and may be attempted again later. The disruption budget task page is blunt about the outcome: if you try to drain a node where an unevictable pod is running, the drain never completes. Node upgrades, repairs and scale-downs all queue behind it.

A budget reaches zero allowed disruptions by two routes. The spec may demand every pod, for example minAvailable equal to the replica count or 100%. Or the spec is reasonable but some pods are unhealthy, so the number of healthy pods is already at or below status.desiredHealthy. Either way, the disruption controller sets the budget’s DisruptionAllowed condition to False with reason InsufficientPods, meaning the pod count is at or below what the budget requires.

Reading the allowance yourself

Terminal window
kubectl get poddisruptionbudgets -A -o json | jq -r '
.items[]
| select(.status.disruptionsAllowed == 0 and .status.expectedPods > 0)
| "\(.metadata.namespace)/\(.metadata.name) healthy=\(.status.currentHealthy) desired=\(.status.desiredHealthy) expected=\(.status.expectedPods)"'

When currentHealthy is below expectedPods, look at the pods before the budget.

Two status numbers, compared as they stand

ZopNight reads the allowed-disruption count and the expected pod count from the budget’s status as last collected. It fires when the allowance is exactly 0 and the budget counts at least one pod. If either number is missing, it stays silent. Because it reads the live allowance rather than the spec, the result reflects both over-strict budgets and budgets drained of headroom by failing pods, and it clears as soon as an evaluation sees room for one eviction.

Budgets handled elsewhere

A budget that counts no pods at all is reported by PDB matches no pods. A spec with maxUnavailable of 0 is also reported by PDB maxUnavailable is zero, so one budget can carry both findings. Keep in mind that node-pressure evictions by the kubelet do not respect budgets at all; a PDB only governs voluntary disruptions.

Stalled upgrades instead of a dollar figure

The finding is a high-severity reliability item with no savings attached. The cost shows up as maintenance windows that overrun, nodes that cannot be retired, and security patches that wait.

Giving the budget room to move

  1. Run kubectl describe poddisruptionbudget <name> -n <namespace> and compare healthy pods with the desired count.
  2. If pods are unhealthy, fix them first; the budget will open up once they report Ready.
  3. If the spec itself demands every pod, lower minAvailable below the replica count or switch to maxUnavailable: 1. A budget may set only one of the two fields.
  4. For workloads with a single replica, add a second one before relying on a budget, as Single replica deployment explains.
  5. Consider unhealthyPodEvictionPolicy: AlwaysAllow, which lets running pods that are not yet healthy be evicted regardless of the budget, so a crash-looping pod cannot hold a drain hostage.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·