Deployments, StatefulSets and DaemonSets with a container that has no liveness probe
What does ZopNight detect here?
ZopNight flags Deployments, StatefulSets and DaemonSets on EKS, GKE and AKS when any main container has no `livenessProbe`. The kubelet treats a missing probe as always passing, so a container that deadlocks but keeps its process alive is never restarted, and it goes on holding its place while doing no useful work.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1707 · RC-1807 · RC-1907 · RC-1708 · RC-1808 · RC-1908 · RC-1718 · RC-1818 · RC-1918 |
| Category | reliability |
| Severity | medium |
| Metric | none — pure configuration read |
| Threshold | any main container without a livenessProbe |
| Source | ZopNight |
| Permissions used | list deployments.apps · list statefulsets.apps · list daemonsets.apps |
Where it applies
A hung process looks healthy to Kubernetes
Kubernetes restarts containers that exit. It cannot restart one that stays running but stops working. That is the gap a liveness probe fills: the probes documentation says liveness probes determine when to restart a container and gives a deadlock as the example, where the application is running but unable to make progress. When the probe fails more times than its tolerance, the kubelet restarts the container.
Without one, the result of that check is fixed. The same page states that if a container does not
provide a particular probe, the kubelet always considers the result a success. A deadlocked
worker then sits in Running indefinitely, holding its CPU and memory reservation, until a person
notices and deletes the pod.
Finding containers with no liveness probe
kubectl get deployments,statefulsets,daemonsets -A -o json | jq -r ' .items[] | .kind as $k | .metadata as $m | .spec.template.spec.containers[] | select(.livenessProbe == null) | "\($k) \($m.namespace)/\($m.name) \(.name)"'First container without a probe triggers it
ZopNight reads the probe settings of each main container in the workload template. As soon as it finds one container with no liveness probe, it raises the recommendation for the workload. Unlike the resource checks, the finding does not enumerate the containers, so use the command above to see which ones need attention. Deployments, StatefulSets and DaemonSets are each checked, on EKS, GKE and AKS, with no metric or window.
When the finding does not apply
Init containers are excluded because they run once and exit, so a liveness probe has nothing to watch. A workload with no containers listed is skipped. The rule cannot judge whether the process already exits on failure, which the probes page says makes a liveness probe unnecessary. Readiness is a separate question, covered by Missing readiness probe.
No saving, but a slow-burning outage
This medium-severity reliability recommendation has no savings figure. The cost is degraded capacity that nobody sees until users do.
Adding a probe that restarts only when it should
- Expose a cheap health endpoint that fails only when the process cannot recover, not when a dependency is slow.
- Add the probe, for example an
httpGeton/healthzwith a sensibleperiodSecondsand a failure threshold of about 3. - For slow-starting apps, add a
startupProbeor raiseinitialDelaySecondsso the liveness probe does not kill the container during boot. - Load-test after the change.