Skip to main content
reliability · kubernetes

Deployments running one replica with no autoscaler behind them

resource types
1
rule IDs covered
3
severity
medium

What does ZopNight detect here?

ZopNight flags an EKS, GKE or AKS Deployment whose `replicas` is exactly 1, unless an HPA or KEDA ScaledObject targets it or its labels, namespace or name mark it as dev or test. With one pod, a node drain, an eviction or a crash leaves the service with nothing serving traffic until a replacement starts.

Signal and threshold

How ZopNight evaluates Deployments running one replica with no autoscaler behind them.
Field Value
Rule IDsRC-1706 · RC-1806 · RC-1906
Categoryreliability
Severitymedium
Metricspec.replicas
Thresholdexactly 1 replica, no HPA or ScaledObject, not dev or test by label, namespace or name
SourceZopNight
Permissions usedlist deployments.apps · list horizontalpodautoscalers.autoscaling

One pod is a single point of failure

.spec.replicas is optional and, per the Deployments documentation, defaults to 1, so a manifest that never mentions it produces a singleton. Everything that removes a pod then becomes an outage. Draining a node for an upgrade, the kubelet evicting under memory pressure, or the container crashing all leave the Deployment with zero ready pods until the replacement is scheduled, pulls its image and passes its readiness probe.

A PodDisruptionBudget cannot rescue a single replica either. A budget that protects the only pod blocks every drain of its node, and a budget that allows one disruption protects nothing.

Finding the singletons

Terminal window
kubectl get deployments -A -o json | jq -r '
.items[] | select(.spec.replicas == 1)
| "\(.metadata.namespace)/\(.metadata.name)"'

Then rule out the ones an autoscaler manages:

Terminal window
kubectl get hpa -A -o json | jq -r '
.items[] | "\(.metadata.namespace) \(.spec.scaleTargetRef.kind)/\(.spec.scaleTargetRef.name)"'

Exactly one replica, and no autoscaler in charge

ZopNight reads the replica count collected for each Deployment and considers only a count of exactly 1. A Deployment at 0 is parked, which Stopped deployment covers, and 2 or more is already redundant. ZopNight then looks for an HPA or KEDA ScaledObject in the same namespace whose target is this Deployment. If one exists, the autoscaler is choosing the replica count on purpose and the rule stays silent.

How dev and test singletons are recognised

Single replicas are normal outside production, so the rule checks the workload’s labels first. A label keyed env, environment, stage or tier with the value prod, production, prd or live keeps the Deployment in scope even if its name looks like a test. The values dev, development, test, testing, qa, staging, stage, sandbox or demo take it out.

With no such label, the namespace and then the name decide. Each is split on hyphens, underscores and dots. A namespace segment of dev, development, test, testing, qa, staging, stage, sandbox or demo suppresses the finding; so does a name segment of any of those or canary, preview or experiment. Whole segments only: contest-api is not treated as a test workload.

The fix adds cost

There is no saving here; the remedy costs money. A second replica doubles the pod’s CPU and memory requests, which may mean more node capacity. The return is availability during upgrades and failures.

Adding redundancy

  1. Scale the Deployment: kubectl scale --replicas=2 deployment/<name> -n <namespace>.
  2. Put the same number in the manifest too. Applying a manifest overwrites a manual scale.
  3. Spread the replicas across nodes or zones, as Workload missing topology spread describes, so both do not share one failure domain.
  4. Add a PodDisruptionBudget with maxUnavailable: 1 so drains move one pod at a time.
  5. Or attach an HPA with minReplicas of 2 if load varies.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·