Failed Kubernetes Jobs, at least an hour old, still on the cluster
What does ZopNight detect here?
ZopNight flags a Kubernetes Job on EKS, GKE or AKS that has no active pods, at least one failed pod and fewer successes than it needs, once the Job is at least 1 hour old. Jobs fail after exhausting `backoffLimit`, 6 retries by default, and Kubernetes keeps the Job and its pods until someone deletes them.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1735 · RC-1835 · RC-1935 |
| Category | reliability |
| Severity | medium |
| Metric | Job status |
| Threshold | Failed, Job at least 1h old |
| Evaluation window | 1h |
| Source | ZopNight |
| Permissions used | list jobs.batch · list pods |
Where it applies
When a Job gives up
A Job retries failing pods until it succeeds or runs out of attempts. The
Jobs documentation sets
.spec.backoffLimit to 6 by default; once that limit is reached the Job is marked failed and its
running pods are terminated. .spec.activeDeadlineSeconds ends a Job the same way when it runs too
long, and the Job’s status becomes type: Failed.
A failed Job means a piece of work did not happen: a data load, a report, a migration, a backup. The same page notes that when a Job finishes, its pods are usually not deleted, so their logs stay available, and the Job object stays too; it is up to the user to delete old Jobs. Left alone, failed Jobs pile up and hide the one that matters.
Listing failed Jobs
kubectl get jobs -A -o json | jq -r ' .items[] | select(any(.status.conditions[]?; .type == "Failed" and .status == "True")) | "\(.metadata.namespace)/\(.metadata.name) created=\(.metadata.creationTimestamp) owner=\(.metadata.ownerReferences[0].name // "none")"'The owner column shows the CronJob, if any, that created each Job.
Failed and at least an hour old
ZopNight treats a Job as failed when it has no active pods, at least one failed pod and fewer successful pods than its required completions. It fires when that is true and the Job was created at least 1 hour ago, going by its creation timestamp. The hour gives a new Job time to work through its retries; because it is measured from creation, not from the failure, a long-running Job that fails late is reported at the next evaluation. When the creation time is missing, the rule does not raise a finding. The finding clears once the Job is deleted, including by a TTL.
Jobs it does not report
Jobs with an active pod and succeeded Jobs are out of scope. Because the test uses pod counts
rather than the Job’s Failed condition, a Job waiting out a retry backoff with no pod running
can read as failed for a short time. The rule does not distinguish Jobs created
by a CronJob from standalone ones, so the failed Job a CronJob keeps as history is flagged too;
CronJobs keep one failed Job by default through failedJobsHistoryLimit. CronJobs that never
create a Job are reported by
CronJob never scheduled.
A reliability finding, not a saving
The recommendation has no dollar figure. Completed pods no longer use CPU or memory; the risk is the work that silently did not happen.
Investigating and cleaning up
- Read the failure reason with
kubectl describe job <name>: backoff limit reached, deadline exceeded, or a pod failure policy. - Read the pod output with
kubectl logs job/<name>, and fix the image, command, permissions or resources behind it. - Re-run it if the work is still needed, from a CronJob with
kubectl create job <name>-rerun --from=cronjob/<cronjob>or by reapplying the manifest. - Delete the failed Job with
kubectl delete job <name>. - Set
.spec.ttlSecondsAfterFinishedon future Jobs so finished Jobs and their pods are removed automatically.