Skip to main content
reliability · kubernetes

StatefulSets with fewer ready pods than desired once the update has finished

resource types
1
rule IDs covered
3
severity
high

What does ZopNight detect here?

ZopNight flags an EKS, GKE or AKS StatefulSet when `readyReplicas` is below `replicas` and `currentReplicas` shows the update already covers every ordinal. Under the default `OrderedReady` policy pods start in order from 0 and each waits for its predecessors to be Running and Ready, so one stuck pod holds back every replica numbered above it.

Signal and threshold

How ZopNight evaluates StatefulSets with fewer ready pods than desired once the update has finished.
Field Value
Rule IDsRC-1715 · RC-1815 · RC-1915
Categoryreliability
Severityhigh
Metricstatus.readyReplicas vs replicas
Thresholdready below desired, update complete
SourceZopNight
Permissions usedlist statefulsets.apps · list pods · list persistentvolumeclaims

Ordered startup turns one bad pod into many missing ones

StatefulSets create pods one at a time. The StatefulSets concept page states the guarantee: for N replicas, pods are created in order from 0 to N-1, and before any pod is scaled, all of its predecessors must be Running and Ready. If web-1 never becomes ready, web-2 and beyond are never created, so a single fault removes more capacity than it would in a Deployment.

Recovery is also less forgiving. The same page warns that with the default OrderedReady policy, a template that never becomes Running and Ready can leave the StatefulSet in a broken state. Reverting the template is not enough: you must also delete the pods that were already started with the bad configuration before the controller retries.

Comparing ready pods with desired

Terminal window
kubectl get statefulsets -A -o json | jq -r '
.items[]
| select(.spec.replicas > 0
and (.status.readyReplicas // 0) < .spec.replicas
and (.status.currentReplicas // .spec.replicas) >= .spec.replicas)
| "\(.metadata.namespace)/\(.metadata.name) ready=\(.status.readyReplicas // 0) desired=\(.spec.replicas)"'

kubectl rollout status statefulset/<name> shows whether Kubernetes still considers an update in progress.

What has to line up before it fires

ZopNight takes the desired replica count and the ready count from the collected StatefulSet. The desired count must be above zero, and a missing ready count is read as zero. Two situations hold the finding back:

  • The current-revision pod count is below desired, meaning pods are still being replaced ordinal by ordinal and a gap is expected.
  • ZopNight recorded a start, stop or scale of this StatefulSet in the last 15 minutes, for instance from a schedule, and ordered startup needs time to finish.

When the current-revision count was not collected at all, the ready gap alone is enough.

StatefulSets it does not judge

A StatefulSet scaled to zero is out of scope. Deployments have their own counterpart, Deployment not fully ready. The finding states the count, not the reason, and one common reason has its own page: Unbound PVC, because a pod whose claim cannot bind never starts.

Lost replicas, no saving

This high-severity reliability finding has no dollar amount. Stateful services usually run a fixed number of members for quorum or replication, so each missing pod erodes fault tolerance.

Unsticking the StatefulSet

  1. Find the lowest-numbered pod that is not ready; the ones above it are usually waiting on it.
  2. Run kubectl describe pod <name>-<ordinal> and read its events for scheduling, image or probe failures.
  3. Check its PersistentVolumeClaim with kubectl get pvc -n <namespace>; a Pending claim blocks the pod.
  4. Read kubectl logs <pod> --previous if the container is restarting.
  5. If a bad template caused it, revert the template and then delete the pods created from the broken revision so the controller rebuilds them.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·