# EKS Cluster Liveness Probe Review Required

> Explains why ZopNight stays silent on EKS liveness probes and how to audit probe coverage with kubectl.

Source: https://zop.dev/integrations/aws/recommendations/eks-cluster-liveness-probe-review-required

---

## What a liveness probe protects against

A liveness probe lets the kubelet notice that a container is stuck and restart it. The
[Kubernetes probe guide](https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/)
describes applications that run for long periods, drift into a broken state such as a deadlock, and
cannot recover except by being restarted. Without a probe, the kubelet has no signal to restart such a container, so it stays in its broken state.

## Why this rule is switched off

The check was designed at cluster level, but probe settings live on each container in each
workload. ZopNight's EKS cluster record has no view of workload specs, and nothing it collects for
the cluster says which containers have probes. The earlier version therefore flagged every EKS cluster
unconditionally, which told you nothing. It now raises no finding at all, and it will stay that way
until per-workload probe data is available.

Other workload checks, such as
<a href="https://zop.dev/integrations/kubernetes/recommendations/missing-cpu-memory-requests">Missing CPU/Memory Requests</a>,
read workload specs directly and are unaffected.

## Auditing probe coverage yourself

List every Deployment container with no liveness probe:

```bash
kubectl get deployments -A -o json | jq -r '
  .items[] | .metadata as $m
  | .spec.template.spec.containers[]
  | select(.livenessProbe == null)
  | "\($m.namespace)/\($m.name) \(.name)"'
```

Run the same query against `statefulsets` and `daemonsets`. Sidecars and short-lived jobs often do
not need a probe, so treat the output as a review list.

## Choosing a probe that helps

A liveness probe should detect "stuck", not "busy". Common forms are `httpGet` against a lightweight
health endpoint, `tcpSocket` for services without HTTP, and `exec` for a command inside the container.
Point it at something that fails only when a restart would help.

## No saving either way

There was never a dollar figure on this finding. The benefit of good probes is fewer silent
failures, not a lower bill.

## Adding liveness probes

1. Add a `livenessProbe` block to each long-running container that can hang.
2. Set `initialDelaySeconds` and `periodSeconds` from real startup time; the Kubernetes example uses
   5 seconds for each. Consider a startup probe for slow-starting apps.
3. Roll out in a non-production namespace and watch restart counts with `kubectl get pods`.
