Skip to main content
rightsizing · kubernetes

Deployments and StatefulSets whose CPU requests sit above their observed peak

resource types
2
rule IDs covered
6
severity
medium

What does ZopNight detect here?

ZopNight currently raises no CPU over-provisioning finding for EKS, GKE or AKS Deployments and StatefulSets, because it needs a per-pod CPU history that metrics-server does not provide. Once that exists, the check compares each per-pod CPU request with the 30-day peak of `CPUUsageMillicores` and prices the reclaimable share between 5% and 85%.

Signal and threshold

How ZopNight evaluates Deployments and StatefulSets whose CPU requests sit above their observed peak.
Field Value
Rule IDsRC-1770 · RC-1870 · RC-1970 · RC-1771 · RC-1871 · RC-1971
Categoryrightsizing
Severitymedium
MetricCPUUsageMillicores
Thresholdpeak usage 5% to 85% below the per-pod CPU request
Evaluation window30d
SourceZopNight
Permissions usedlist deployments.apps · list statefulsets.apps · get pods.metrics.k8s.io

Requested CPU is paid for whether it is used or not

The scheduler places pods by their requests, not their usage. The resource management page says the scheduler ensures the sum of requests on a node stays below its capacity, and refuses to place a pod when that check fails even if actual CPU usage on the node is very low. A request of 2000 millicores for a pod that peaks at 400 therefore holds 1600 millicores of node capacity that nothing else may schedule into, and enough of those gaps add whole nodes to the bill.

CPU is also the safer resource to trim. The kernel enforces CPU limits by throttling rather than killing, so a container that ends up short of CPU slows down instead of restarting.

Comparing requests with real usage yourself

List each workload’s per-container CPU requests:

Terminal window
kubectl get deployments,statefulsets -A -o json | jq -r '
.items[] | .kind as $k | .metadata as $m
| "\($k) \($m.namespace)/\($m.name) "
+ ([.spec.template.spec.containers[] | "\(.name)=\(.resources.requests.cpu // "none")"] | join(" "))'

kubectl top pod --containers -n <namespace> gives current usage, but a single reading is not a peak. For a history, run the Vertical Pod Autoscaler with updateMode: "Off", which records recommendations in the VPA’s status without touching the pods.

Evidence ZopNight needs before it sizes down

Every one of these must be present, or the rule raises nothing:

  • The workload’s collected status is running. Any other status is skipped.
  • Its containers declare CPU requests, summed per pod, above zero.
  • ZopNight has a monthly cost attributed to the workload’s pods.
  • A per-pod CPU usage series spans the 30-day window with at least 7 days of coverage. A single point-in-time sample does not qualify. ZopNight’s usage source today, metrics-server, gives only point-in-time readings, so this check stays silent until a windowed per-pod CPU history is available.
  • The observed peak is below the request, and the gap is between 5% and 85% of the request. Below 5% the change is noise; above 85% the reading or the request is more likely wrong than the workload idle.

How the saving is calculated

Terminal window
reclaimable share = (per-pod CPU request - peak CPU usage) / per-pod CPU request
monthly saving = attributed monthly cost x reclaimable share, capped at the cost

The peak is the highest usage observed, not an average, so the suggested request still covers the busiest moment in the window. Workloads with no CPU request at all are reported by Missing CPU/memory requests instead, and memory headroom by Memory over-provisioned.

Lowering the request safely

  1. Compare the finding’s peak with VPA’s target recommendation if you run one.
  2. Set the new request at the peak plus a margin, for example with kubectl set resources deployment <name> -c=<container> --requests=cpu=500m.
  3. If an HPA scales the workload on CPU utilization, remember utilization is measured as a share of the request, so a smaller request raises the reported figure and can add replicas. Adjust the HPA target in the same change.
  4. Roll out and watch latency and CPU throttling before trimming further.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·