Skip to main content
governance · kubernetes

CronJobs left on concurrencyPolicy Allow, where slow runs can overlap

resource types
1
rule IDs covered
3
severity
low

What does ZopNight detect here?

ZopNight flags any Kubernetes CronJob on EKS, GKE or AKS whose `concurrencyPolicy` is `Allow`. `Allow` is the Kubernetes default, so every CronJob that never chose `Forbid` or `Replace` qualifies: when one run outlasts its schedule interval, the next Job starts beside it and the overlapping pods compete for data and node capacity.

Signal and threshold

How ZopNight evaluates CronJobs left on concurrencyPolicy Allow, where slow runs can overlap.
Field Value
Rule IDsRC-1750 · RC-1850 · RC-1950
Categorygovernance
Severitylow
Metricspec.concurrencyPolicy
Thresholdequals Allow
SourceZopNight
Permissions usedlist cronjobs.batch

What Allow does when a run is slow

A CronJob creates a Job each time its schedule fires. The CronJob concepts page lists three concurrency policies and marks Allow as the default: concurrently running Jobs are permitted. Forbid skips the new run if the previous one is still going, and Replace cancels the running Job and starts the new one in its place.

Under Allow, a job scheduled every five minutes that suddenly takes twelve leaves two or three copies running at once. If the job writes to a database, processes a queue or rotates files, the copies race each other. If it is heavy, each copy requests its own CPU and memory, so a backlog of overlapping runs can crowd other pods off the nodes. The same page also warns that scheduling is approximate and that jobs should be idempotent, because Kubernetes can create two Jobs for one slot in rare cases.

Finding CronJobs that allow overlap

Kubernetes fills in the default when the object is stored, so the field is almost always present. This lists every CronJob still set to Allow, with its schedule:

Terminal window
kubectl get cronjobs -A -o json | jq -r '
.items[]
| select((.spec.concurrencyPolicy // "Allow") == "Allow")
| "\(.metadata.namespace)/\(.metadata.name) schedule=\(.spec.schedule)"'

Compare that list with job durations from kubectl get jobs -A to find the ones at real risk.

One field decides the finding

The check reads concurrencyPolicy as ZopNight collected it and fires only when the value is exactly Allow. There is no threshold on job duration and no history window: ZopNight does not wait to see an overlap happen before flagging the setting. When the policy is not reported at all, the CronJob is skipped.

Cases it does not try to judge

The rule cannot tell whether overlap is harmless for a given job. A job that always finishes in seconds on an hourly schedule will be flagged even though it will never overlap in practice, and some batch designs rely on parallel runs on purpose. Suspended CronJobs are not excluded; they are also reported by Suspended CronJob.

Governance finding, no dollar figure

This is a low-severity governance recommendation with no savings estimate. The cost it guards against is indirect: duplicated work, corrupted output and the extra node capacity that stacked runs consume.

Choosing Forbid or Replace

  1. Decide which failure mode is acceptable. Forbid keeps the running job and skips the new slot; Replace kills the running job so the newest run wins.
  2. Patch the CronJob, for example kubectl patch cronjob nightly-report -p '{"spec":{"concurrencyPolicy":"Forbid"}}'.
  3. With Forbid, set startingDeadlineSeconds deliberately. Kubernetes counts skipped slots as missed, and when it finds more than 100 missed schedules it does not start the catch-up Job and logs too many missed start times. A deadline limits how far back that count looks.
  4. Add activeDeadlineSeconds to the Job template so a hung run cannot hold the slot forever.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·