Deployments and StatefulSets with memory requests above their working set
What does ZopNight detect here?
ZopNight currently raises no memory over-provisioning finding for EKS, GKE or AKS Deployments and StatefulSets. Memory is enforced by OOM kills, so a request cut too far crashes pods, and an HPA `ScalingLimited` replica signal says nothing about per-pod memory. The check stays silent until a measured per-pod memory working set is available.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1772 · RC-1872 · RC-1972 · RC-1773 · RC-1873 · RC-1973 |
| Category | rightsizing |
| Severity | medium |
| Metric | none — pure configuration read |
| Source | ZopNight |
| Permissions used | list deployments.apps · list statefulsets.apps · list horizontalpodautoscalers.autoscaling · get pods.metrics.k8s.io |
Where it applies
Oversized memory requests strand node capacity
Like CPU, memory requests reserve room on a node whether or not the container uses it. The scheduler, as the resource management page puts it, refuses to place a pod when the requests on a node would exceed capacity, even if actual usage there is very low. A pod requesting 4Gi that settles at 1.5Gi keeps 2.5Gi idle on every node that runs a copy of it, and that idle memory is capacity you are paying for.
Why memory cuts need stronger evidence than CPU
Memory does not degrade gracefully. The same page explains that memory limits are enforced by the kernel with OOM kills, reactively, when it detects memory pressure. A request set below the real working set also makes the pod an early eviction candidate, since the kubelet evicts pods using more than they requested first when a node runs short. An aggressive memory recommendation can therefore turn a saving into restarts, and for StatefulSets into disrupted replicas.
A replica signal is not a memory measurement
An HPA reports ScalingLimited when the scale it calculated had to be capped at its bounds; with
the reason TooFewReplicas it wanted fewer pods than minReplicas allows. That describes the
replica count only. A workload can have too many replicas while each pod’s memory request is
correct, or even too small. ZopNight does not treat that condition as evidence that memory
requests are oversized. It has no per-pod memory working-set measurement to use instead, so
today this check raises no finding on any workload rather than guess.
Measuring memory headroom yourself
List memory requests per container:
kubectl get deployments,statefulsets -A -o json | jq -r ' .items[] | .kind as $k | .metadata as $m | "\($k) \($m.namespace)/\($m.name) " + ([.spec.template.spec.containers[] | "\(.name)=\(.resources.requests.memory // "none")"] | join(" "))'Compare with usage over time, not a single kubectl top pod --containers reading. The Vertical
Pod Autoscaler in updateMode: "Off" stores target, lower-bound and upper-bound recommendations
in its status without changing the pods, which gives a defensible basis.
No estimate without a measurement
ZopNight does not put a dollar figure on memory headroom it has not measured. CPU headroom, which is safer to reclaim, is covered by CPU over-provisioned.
Trimming memory without OOM kills
- Start from the observed peak working set, not the average, and add a margin.
- Lower the request in small steps, and keep the memory limit at or above the new request.
- Watch container restarts and eviction events for several days before the next step.
- For StatefulSets, change one replica at a time where the application allows it.
- On clusters with in-place pod resize, set the container’s
resizePolicyfor memory toRestartContainerif the application cannot adopt a new limit while running.