Skip to main content
reliability · kubernetes

Production Deployments and StatefulSets with several replicas but no topology spread constraints

resource types
2
rule IDs covered
6
severity
medium

What does ZopNight detect here?

ZopNight flags a production-tagged Deployment or StatefulSet on EKS, GKE or AKS that runs 2 or more replicas without any `topologySpreadConstraints`. The scheduler's built-in defaults only nudge placement, with a `maxSkew` of 3 per node and 5 per zone under `ScheduleAnyway`, so several replicas can still share one node or zone.

Signal and threshold

How ZopNight evaluates Production Deployments and StatefulSets with several replicas but no topology spread constraints.
Field Value
Rule IDsRC-1757 · RC-1857 · RC-1957 · RC-1758 · RC-1858 · RC-1958
Categoryreliability
Severitymedium
MetrictopologySpreadConstraints
Thresholdnone declared, 2 or more replicas, production tag
SourceZopNight
Permissions usedlist deployments.apps · list statefulsets.apps

Replicas are only redundant if they are apart

Running three replicas protects nothing if all three land on the same node or in the same zone. Pod topology spread constraints let you tell the scheduler how evenly to distribute matching pods across failure domains such as nodes (kubernetes.io/hostname) and zones (topology.kubernetes.io/zone). maxSkew caps the difference in pod count between domains, and whenUnsatisfiable chooses between refusing to schedule (DoNotSchedule) and merely preferring balance (ScheduleAnyway).

Without constraints of your own, and with no cluster-level defaults configured, kube-scheduler applies built-in defaults: maxSkew 3 across hostnames and 5 across zones, both ScheduleAnyway. Those are soft preferences with generous skew. A five-replica service can legally pile most of its pods into one zone, and a node or zone failure then removes most of its capacity at once.

Checking which workloads declare constraints

Terminal window
kubectl get deployments,statefulsets -A -o json | jq -r '
.items[]
| select(.spec.replicas >= 2
and ((.spec.template.spec.topologySpreadConstraints // []) | length) == 0)
| "\(.kind) \(.metadata.namespace)/\(.metadata.name) replicas=\(.spec.replicas)"'

To see where the pods of one workload actually run:

Terminal window
kubectl get pods -n <namespace> -l <selector> -o wide

Production tag, two replicas, no constraints

ZopNight applies this check to Deployments and StatefulSets on EKS, GKE and AKS, and all of the following must hold:

  • The workload carries a production environment tag: a key of env, environment, stage or tier with the value prod, production, prd or live.
  • Its replica count is at least 2, since spreading a single pod is meaningless.
  • Its pod template lists no topologySpreadConstraints.

Workloads outside the check

Anything without a production tag is skipped, whatever its name. Single-replica workloads are the concern of Single replica deployment instead. Other placement tools, such as pod anti-affinity or node selectors, are not taken into account, which is the main source of false positives.

Correlated failure, not a saving

The recommendation has no savings figure. The exposure is a single node or zone outage taking down every replica of a production service at once.

Adding spread constraints

  1. Add a constraint on topology.kubernetes.io/zone to the pod template, with a labelSelector matching the workload’s own pod labels and maxSkew: 1.
  2. Add a second constraint on kubernetes.io/hostname so replicas also avoid sharing a node.
  3. Choose whenUnsatisfiable deliberately. DoNotSchedule guarantees the spread but can leave pods pending when a zone lacks capacity; ScheduleAnyway never blocks but only prefers.
  4. Roll out and check placement with kubectl get pods -o wide.
  5. Constraints are not re-checked when pods are removed, so a scale-down can leave the spread uneven. A tool such as the Descheduler can rebalance it.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·