Production Deployments and StatefulSets with several replicas but no topology spread constraints
What does ZopNight detect here?
ZopNight flags a production-tagged Deployment or StatefulSet on EKS, GKE or AKS that runs 2 or more replicas without any `topologySpreadConstraints`. The scheduler's built-in defaults only nudge placement, with a `maxSkew` of 3 per node and 5 per zone under `ScheduleAnyway`, so several replicas can still share one node or zone.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1757 · RC-1857 · RC-1957 · RC-1758 · RC-1858 · RC-1958 |
| Category | reliability |
| Severity | medium |
| Metric | topologySpreadConstraints |
| Threshold | none declared, 2 or more replicas, production tag |
| Source | ZopNight |
| Permissions used | list deployments.apps · list statefulsets.apps |
Where it applies
Replicas are only redundant if they are apart
Running three replicas protects nothing if all three land on the same node or in the same zone.
Pod topology spread constraints
let you tell the scheduler how evenly to distribute matching pods across failure domains such as
nodes (kubernetes.io/hostname) and zones (topology.kubernetes.io/zone). maxSkew caps the
difference in pod count between domains, and whenUnsatisfiable chooses between refusing to
schedule (DoNotSchedule) and merely preferring balance (ScheduleAnyway).
Without constraints of your own, and with no cluster-level defaults configured, kube-scheduler
applies built-in defaults: maxSkew 3 across hostnames and 5 across zones, both ScheduleAnyway.
Those are soft preferences with generous skew. A five-replica service can legally pile most of its
pods into one zone, and a node or zone failure then removes most of its capacity at once.
Checking which workloads declare constraints
kubectl get deployments,statefulsets -A -o json | jq -r ' .items[] | select(.spec.replicas >= 2 and ((.spec.template.spec.topologySpreadConstraints // []) | length) == 0) | "\(.kind) \(.metadata.namespace)/\(.metadata.name) replicas=\(.spec.replicas)"'To see where the pods of one workload actually run:
kubectl get pods -n <namespace> -l <selector> -o wideProduction tag, two replicas, no constraints
ZopNight applies this check to Deployments and StatefulSets on EKS, GKE and AKS, and all of the following must hold:
- The workload carries a production environment tag: a key of
env,environment,stageortierwith the valueprod,production,prdorlive. - Its replica count is at least 2, since spreading a single pod is meaningless.
- Its pod template lists no
topologySpreadConstraints.
Workloads outside the check
Anything without a production tag is skipped, whatever its name. Single-replica workloads are the concern of Single replica deployment instead. Other placement tools, such as pod anti-affinity or node selectors, are not taken into account, which is the main source of false positives.
Correlated failure, not a saving
The recommendation has no savings figure. The exposure is a single node or zone outage taking down every replica of a production service at once.
Adding spread constraints
- Add a constraint on
topology.kubernetes.io/zoneto the pod template, with alabelSelectormatching the workload’s own pod labels andmaxSkew: 1. - Add a second constraint on
kubernetes.io/hostnameso replicas also avoid sharing a node. - Choose
whenUnsatisfiabledeliberately.DoNotScheduleguarantees the spread but can leave pods pending when a zone lacks capacity;ScheduleAnywaynever blocks but only prefers. - Roll out and check placement with
kubectl get pods -o wide. - Constraints are not re-checked when pods are removed, so a scale-down can leave the spread uneven. A tool such as the Descheduler can rebalance it.