EKS liveness probe review: why ZopNight no longer raises this at cluster level
What does ZopNight detect here?
ZopNight does not currently raise this EKS finding. An earlier version flagged every cluster without evidence, because checking `livenessProbe` settings requires reading each workload through the Kubernetes API. Until per-workload probe data is collected, the rule stays silent; use `kubectl get deployments -A -o json` to find containers without a liveness probe yourself.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-063 |
| Category | compliance |
| Severity | medium |
| Metric | none — pure configuration read |
| Source | ZopNight |
| Permissions used | eks:DescribeCluster |
Where it applies
What a liveness probe protects against
A liveness probe lets the kubelet notice that a container is stuck and restart it. The Kubernetes probe guide describes applications that run for long periods, drift into a broken state such as a deadlock, and cannot recover except by being restarted. Without a probe, the kubelet has no signal to restart such a container, so it stays in its broken state.
Why this rule is switched off
The check was designed at cluster level, but probe settings live on each container in each workload. ZopNight’s EKS cluster record has no view of workload specs, and nothing it collects for the cluster says which containers have probes. The earlier version therefore flagged every EKS cluster unconditionally, which told you nothing. It now raises no finding at all, and it will stay that way until per-workload probe data is available.
Other workload checks, such as Missing CPU/Memory Requests, read workload specs directly and are unaffected.
Auditing probe coverage yourself
List every Deployment container with no liveness probe:
kubectl get deployments -A -o json | jq -r ' .items[] | .metadata as $m | .spec.template.spec.containers[] | select(.livenessProbe == null) | "\($m.namespace)/\($m.name) \(.name)"'Run the same query against statefulsets and daemonsets. Sidecars and short-lived jobs often do
not need a probe, so treat the output as a review list.
Choosing a probe that helps
A liveness probe should detect “stuck”, not “busy”. Common forms are httpGet against a lightweight
health endpoint, tcpSocket for services without HTTP, and exec for a command inside the container.
Point it at something that fails only when a restart would help.
No saving either way
There was never a dollar figure on this finding. The benefit of good probes is fewer silent failures, not a lower bill.
Adding liveness probes
- Add a
livenessProbeblock to each long-running container that can hang. - Set
initialDelaySecondsandperiodSecondsfrom real startup time; the Kubernetes example uses 5 seconds for each. Consider a startup probe for slow-starting apps. - Roll out in a non-production namespace and watch restart counts with
kubectl get pods.