Skip to main content
reliability · kubernetes

Failed Kubernetes Jobs, at least an hour old, still on the cluster

resource types
1
rule IDs covered
3
severity
medium

What does ZopNight detect here?

ZopNight flags a Kubernetes Job on EKS, GKE or AKS that has no active pods, at least one failed pod and fewer successes than it needs, once the Job is at least 1 hour old. Jobs fail after exhausting `backoffLimit`, 6 retries by default, and Kubernetes keeps the Job and its pods until someone deletes them.

Signal and threshold

How ZopNight evaluates Failed Kubernetes Jobs, at least an hour old, still on the cluster.
Field Value
Rule IDsRC-1735 · RC-1835 · RC-1935
Categoryreliability
Severitymedium
MetricJob status
ThresholdFailed, Job at least 1h old
Evaluation window1h
SourceZopNight
Permissions usedlist jobs.batch · list pods

When a Job gives up

A Job retries failing pods until it succeeds or runs out of attempts. The Jobs documentation sets .spec.backoffLimit to 6 by default; once that limit is reached the Job is marked failed and its running pods are terminated. .spec.activeDeadlineSeconds ends a Job the same way when it runs too long, and the Job’s status becomes type: Failed.

A failed Job means a piece of work did not happen: a data load, a report, a migration, a backup. The same page notes that when a Job finishes, its pods are usually not deleted, so their logs stay available, and the Job object stays too; it is up to the user to delete old Jobs. Left alone, failed Jobs pile up and hide the one that matters.

Listing failed Jobs

Terminal window
kubectl get jobs -A -o json | jq -r '
.items[]
| select(any(.status.conditions[]?; .type == "Failed" and .status == "True"))
| "\(.metadata.namespace)/\(.metadata.name) created=\(.metadata.creationTimestamp) owner=\(.metadata.ownerReferences[0].name // "none")"'

The owner column shows the CronJob, if any, that created each Job.

Failed and at least an hour old

ZopNight treats a Job as failed when it has no active pods, at least one failed pod and fewer successful pods than its required completions. It fires when that is true and the Job was created at least 1 hour ago, going by its creation timestamp. The hour gives a new Job time to work through its retries; because it is measured from creation, not from the failure, a long-running Job that fails late is reported at the next evaluation. When the creation time is missing, the rule does not raise a finding. The finding clears once the Job is deleted, including by a TTL.

Jobs it does not report

Jobs with an active pod and succeeded Jobs are out of scope. Because the test uses pod counts rather than the Job’s Failed condition, a Job waiting out a retry backoff with no pod running can read as failed for a short time. The rule does not distinguish Jobs created by a CronJob from standalone ones, so the failed Job a CronJob keeps as history is flagged too; CronJobs keep one failed Job by default through failedJobsHistoryLimit. CronJobs that never create a Job are reported by CronJob never scheduled.

A reliability finding, not a saving

The recommendation has no dollar figure. Completed pods no longer use CPU or memory; the risk is the work that silently did not happen.

Investigating and cleaning up

  1. Read the failure reason with kubectl describe job <name>: backoff limit reached, deadline exceeded, or a pod failure policy.
  2. Read the pod output with kubectl logs job/<name>, and fix the image, command, permissions or resources behind it.
  3. Re-run it if the work is still needed, from a CronJob with kubectl create job <name>-rerun --from=cronjob/<cronjob> or by reapplying the manifest.
  4. Delete the failed Job with kubectl delete job <name>.
  5. Set .spec.ttlSecondsAfterFinished on future Jobs so finished Jobs and their pods are removed automatically.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·