Deployments and StatefulSets whose CPU requests sit above their observed peak
What does ZopNight detect here?
ZopNight currently raises no CPU over-provisioning finding for EKS, GKE or AKS Deployments and StatefulSets, because it needs a per-pod CPU history that metrics-server does not provide. Once that exists, the check compares each per-pod CPU request with the 30-day peak of `CPUUsageMillicores` and prices the reclaimable share between 5% and 85%.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1770 · RC-1870 · RC-1970 · RC-1771 · RC-1871 · RC-1971 |
| Category | rightsizing |
| Severity | medium |
| Metric | CPUUsageMillicores |
| Threshold | peak usage 5% to 85% below the per-pod CPU request |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | list deployments.apps · list statefulsets.apps · get pods.metrics.k8s.io |
Where it applies
Requested CPU is paid for whether it is used or not
The scheduler places pods by their requests, not their usage. The resource management page says the scheduler ensures the sum of requests on a node stays below its capacity, and refuses to place a pod when that check fails even if actual CPU usage on the node is very low. A request of 2000 millicores for a pod that peaks at 400 therefore holds 1600 millicores of node capacity that nothing else may schedule into, and enough of those gaps add whole nodes to the bill.
CPU is also the safer resource to trim. The kernel enforces CPU limits by throttling rather than killing, so a container that ends up short of CPU slows down instead of restarting.
Comparing requests with real usage yourself
List each workload’s per-container CPU requests:
kubectl get deployments,statefulsets -A -o json | jq -r ' .items[] | .kind as $k | .metadata as $m | "\($k) \($m.namespace)/\($m.name) " + ([.spec.template.spec.containers[] | "\(.name)=\(.resources.requests.cpu // "none")"] | join(" "))'kubectl top pod --containers -n <namespace> gives current usage, but a single reading is not a
peak. For a history, run the Vertical Pod Autoscaler with updateMode: "Off", which records
recommendations in the VPA’s status without touching the pods.
Evidence ZopNight needs before it sizes down
Every one of these must be present, or the rule raises nothing:
- The workload’s collected status is running. Any other status is skipped.
- Its containers declare CPU requests, summed per pod, above zero.
- ZopNight has a monthly cost attributed to the workload’s pods.
- A per-pod CPU usage series spans the 30-day window with at least 7 days of coverage. A single point-in-time sample does not qualify. ZopNight’s usage source today, metrics-server, gives only point-in-time readings, so this check stays silent until a windowed per-pod CPU history is available.
- The observed peak is below the request, and the gap is between 5% and 85% of the request. Below 5% the change is noise; above 85% the reading or the request is more likely wrong than the workload idle.
How the saving is calculated
reclaimable share = (per-pod CPU request - peak CPU usage) / per-pod CPU requestmonthly saving = attributed monthly cost x reclaimable share, capped at the costThe peak is the highest usage observed, not an average, so the suggested request still covers the busiest moment in the window. Workloads with no CPU request at all are reported by Missing CPU/memory requests instead, and memory headroom by Memory over-provisioned.
Lowering the request safely
- Compare the finding’s peak with VPA’s target recommendation if you run one.
- Set the new request at the peak plus a margin, for example with
kubectl set resources deployment <name> -c=<container> --requests=cpu=500m. - If an HPA scales the workload on CPU utilization, remember utilization is measured as a share of the request, so a smaller request raises the reported figure and can add replicas. Adjust the HPA target in the same change.
- Roll out and watch latency and CPU throttling before trimming further.