StatefulSets with fewer ready pods than desired once the update has finished
What does ZopNight detect here?
ZopNight flags an EKS, GKE or AKS StatefulSet when `readyReplicas` is below `replicas` and `currentReplicas` shows the update already covers every ordinal. Under the default `OrderedReady` policy pods start in order from 0 and each waits for its predecessors to be Running and Ready, so one stuck pod holds back every replica numbered above it.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1715 · RC-1815 · RC-1915 |
| Category | reliability |
| Severity | high |
| Metric | status.readyReplicas vs replicas |
| Threshold | ready below desired, update complete |
| Source | ZopNight |
| Permissions used | list statefulsets.apps · list pods · list persistentvolumeclaims |
Where it applies
Ordered startup turns one bad pod into many missing ones
StatefulSets create pods one at a time. The
StatefulSets concept page
states the guarantee: for N replicas, pods are created in order from 0 to N-1, and before any pod
is scaled, all of its predecessors must be Running and Ready. If web-1 never becomes ready,
web-2 and beyond are never created, so a single fault removes more capacity than it would in a
Deployment.
Recovery is also less forgiving. The same page warns that with the default OrderedReady
policy, a template that never becomes Running and Ready can leave the StatefulSet in a broken
state. Reverting the template is not enough: you must also delete the pods that were already
started with the bad configuration before the controller retries.
Comparing ready pods with desired
kubectl get statefulsets -A -o json | jq -r ' .items[] | select(.spec.replicas > 0 and (.status.readyReplicas // 0) < .spec.replicas and (.status.currentReplicas // .spec.replicas) >= .spec.replicas) | "\(.metadata.namespace)/\(.metadata.name) ready=\(.status.readyReplicas // 0) desired=\(.spec.replicas)"'kubectl rollout status statefulset/<name> shows whether Kubernetes still considers an update in
progress.
What has to line up before it fires
ZopNight takes the desired replica count and the ready count from the collected StatefulSet. The desired count must be above zero, and a missing ready count is read as zero. Two situations hold the finding back:
- The current-revision pod count is below desired, meaning pods are still being replaced ordinal by ordinal and a gap is expected.
- ZopNight recorded a start, stop or scale of this StatefulSet in the last 15 minutes, for instance from a schedule, and ordered startup needs time to finish.
When the current-revision count was not collected at all, the ready gap alone is enough.
StatefulSets it does not judge
A StatefulSet scaled to zero is out of scope. Deployments have their own counterpart, Deployment not fully ready. The finding states the count, not the reason, and one common reason has its own page: Unbound PVC, because a pod whose claim cannot bind never starts.
Lost replicas, no saving
This high-severity reliability finding has no dollar amount. Stateful services usually run a fixed number of members for quorum or replication, so each missing pod erodes fault tolerance.
Unsticking the StatefulSet
- Find the lowest-numbered pod that is not ready; the ones above it are usually waiting on it.
- Run
kubectl describe pod <name>-<ordinal>and read its events for scheduling, image or probe failures. - Check its PersistentVolumeClaim with
kubectl get pvc -n <namespace>; aPendingclaim blocks the pod. - Read
kubectl logs <pod> --previousif the container is restarting. - If a bad template caused it, revert the template and then delete the pods created from the broken revision so the controller rebuilds them.