Skip to main content
reliability · kubernetes

Deployments using the Recreate strategy, which takes the app down on every rollout

resource types
1
rule IDs covered
3
severity
medium

What does ZopNight detect here?

ZopNight flags a Kubernetes Deployment on EKS, GKE or AKS whose `spec.strategy.type` is `Recreate`. With Recreate, Kubernetes terminates every existing pod before creating any new one, so each rollout is a short outage. The default `RollingUpdate` keeps at least 75% of pods serving throughout an update.

Signal and threshold

How ZopNight evaluates Deployments using the Recreate strategy, which takes the app down on every rollout.
Field Value
Rule IDsRC-1759 · RC-1859 · RC-1959
Categoryreliability
Severitymedium
Metricspec.strategy.type
Thresholdequals Recreate
SourceZopNight
Permissions usedlist deployments.apps

Recreate turns every deploy into downtime

The Deployments documentation defines two strategies. RollingUpdate is the default: by default it keeps at least 75% of the desired pods up (25% max unavailable) and runs at most 125% (25% max surge) while it replaces them. Recreate kills all existing pods before new ones are created, and for an upgrade it waits for their removal to succeed before starting any pod of the new revision.

The gap between the last old pod stopping and the first new pod passing readiness is time with no capacity at all, and it lasts as long as the new pods take to pull their image, start and pass readiness. It also makes a bad release more painful, because there is no old version still serving while you notice.

Listing Deployments on Recreate

Terminal window
kubectl get deployments -A -o json | jq -r '
.items[]
| select(.spec.strategy.type == "Recreate")
| "\(.metadata.namespace)/\(.metadata.name) replicas=\(.spec.replicas)"'

The strategy field is the whole test

ZopNight reads the Deployment’s strategy as collected and fires when it is exactly Recreate. The replica count does not matter: a three-replica Deployment on Recreate has the same outage on each rollout as a single-replica one. There is no metric, threshold or window, and the finding clears once the strategy changes. A Deployment with no reported strategy is skipped.

Legitimate reasons to keep Recreate

The rule cannot see why Recreate was chosen, and sometimes it is the right call. Two versions of an app that write to the same data in incompatible ways, or a singleton that must never have two active copies, may not be safe to overlap. StatefulSets and DaemonSets have their own update strategies and are not part of this check. Single-replica Deployments are covered separately by Single replica deployment, because a rolling update of one replica still leaves little margin.

Availability risk, no saving attached

This medium-severity reliability finding has no savings estimate. The cost is user-facing errors during each release window.

Moving to a rolling update

  1. Confirm the app tolerates two versions running for a short time: shared schemas, queue consumers and file locks are the usual concerns.
  2. Make sure the pods have a readiness probe, so the rollout waits for new pods to be ready before removing old ones.
  3. Patch the strategy: kubectl patch deployment <name> -p '{"spec":{"strategy":{"type":"RollingUpdate"}}}', then tune maxSurge and maxUnavailable if the 25% defaults do not suit the replica count.
  4. Watch the next release with kubectl rollout status deployment/<name>.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·