Skip to main content
reliability · kubernetes

Deployments with fewer ready replicas than desired after the rollout finished

resource types
1
rule IDs covered
3
severity
high

What does ZopNight detect here?

ZopNight flags a Kubernetes Deployment on EKS, GKE or AKS when `readyReplicas` is below the desired replica count while the rollout itself is complete, meaning `updatedReplicas` already equals the desired count. The Deployment is running short-handed with no update to blame, usually because pods are crash-looping, pending or failing readiness checks.

Signal and threshold

How ZopNight evaluates Deployments with fewer ready replicas than desired after the rollout finished.
Field Value
Rule IDsRC-1714 · RC-1814 · RC-1914
Categoryreliability
Severityhigh
Metricstatus.readyReplicas vs desired replicas
Thresholdready below desired, rollout complete
SourceZopNight
Permissions usedlist deployments.apps · list pods

Short on capacity with nothing in progress

During a rollout it is normal for ready replicas to trail desired while new pods start. Once every replica has been updated, the two numbers should match. When they do not, some pods are stuck, and the service is running on less capacity than it was sized for. With the rolling update defaults in the Deployments documentation, up to 25% of pods may be unavailable during an update; a Deployment stuck below desired afterwards has lost that margin for good.

The same page lists the usual reasons a Deployment stops progressing: insufficient quota, failing readiness probes, image pull errors, insufficient permissions, limit ranges and application misconfiguration. Kubernetes takes no action on a stalled Deployment beyond reporting a Progressing condition with reason ProgressDeadlineExceeded.

Finding Deployments running below desired

Terminal window
kubectl get deployments -A -o json | jq -r '
.items[]
| select(.spec.replicas > 0
and (.status.readyReplicas // 0) < .spec.replicas
and (.status.updatedReplicas // 0) >= .spec.replicas)
| "\(.metadata.namespace)/\(.metadata.name) ready=\(.status.readyReplicas // 0) desired=\(.spec.replicas)"'

kubectl rollout status deployment/<name> confirms whether Kubernetes considers the rollout finished.

Three checks before the finding fires

ZopNight reads the desired and ready replica counts and fires when desired is above zero and ready is below it. A missing ready count is treated as zero. Two conditions hold it back:

  • If the updated replica count is below desired, a rollout is still running and the gap is expected.
  • If ZopNight itself started, stopped or scaled the Deployment within the last 15 minutes, for example through a schedule, it waits for pods to settle.

There is no longer window; the finding reflects the counts at evaluation time.

Deployments outside the check

A Deployment set to zero replicas has nothing to be ready and is skipped; parked Deployments are covered by Stopped deployment. StatefulSets have their own version of this check, StatefulSet not fully ready. The finding reports the counts, not the cause, so the diagnosis steps below are still needed.

Reliability risk without a saving

This high-severity reliability finding carries no dollar amount. The exposure is reduced capacity and redundancy for a service that believes it has its full replica count.

Diagnosing the missing replicas

  1. Run kubectl describe deployment <name> and read the conditions for ProgressDeadlineExceeded or quota errors.
  2. List the pods and find those not ready: kubectl get pods -n <namespace> -l <selector>.
  3. For Pending pods, read kubectl describe pod events for scheduling failures such as insufficient CPU or unbound volumes.
  4. For CrashLoopBackOff or failing readiness, read kubectl logs <pod> --previous and the probe configuration.
  5. Fix the cause and confirm ready replicas return to the desired count.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·