GKE node pools whose node CPU peaked under 20% and memory under 30% for 30 days
What does ZopNight detect here?
ZopNight's check targets GKE node pools whose peak node CPU stayed under 20% and peak memory under 30%, and would price a move to the next-smaller machine type. It raises no findings today: node VMs are costed individually, so a pool carries no cost of its own, and node metrics are not yet matched to their pool.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-121 |
| Category | rightsizing |
| Severity | medium |
| Metric | node CPU and memory utilization |
| Threshold | peak CPU under 20% and peak memory under 30% |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | container.clusters.get · container.nodes.list · monitoring.timeSeries.list |
Where it applies
Nodes that stay nearly empty
The cluster autoscaler removes nodes when they are underutilized and their pods can be scheduled elsewhere, but only on pools where it is enabled and only down to the pool’s minimum. A pool with a high minimum node count, or with autoscaling off, keeps every node whether or not pods need it.
Low peaks, not low averages, are the telling sign. If a pool’s busiest moment in a month still leaves four-fifths of its CPU idle, the capacity is not headroom; it is surplus.
Checking node usage in a pool
gcloud container node-pools list --cluster=CLUSTER_NAME --location=LOCATIONIn Metrics Explorer, chart the maximum of kubernetes.io/node/cpu/allocatable_utilization and
kubernetes.io/node/memory/allocatable_utilization over 30 days, filtered to the pool’s nodes.
Both peaks must be low
ZopNight is built to read 30 days of per-node CPU and memory for the pool. It would fire only when peak CPU stayed under 20% and peak memory under 30%. Requiring both keeps memory-bound pools with idle CPU off the list, which makes this a stricter test than the CPU-only check behind GKE Node Pool Machine Type Rightsizing.
Where no finding appears
Today no pool gets a finding: GKE node VMs are costed one by one, so a pool has no cost of its own, and node metrics are not yet matched to their pool. Beyond that, no utilization data for the pool means no finding. The pool must also have a smaller same-family machine type to move to and known rates for both types; custom types and pools at the family floor are not scored.
Downsizing the machine type
saving = current pool cost x (current type rate - smaller type rate) / current type rateRemoving whole nodes can save more than this figure, which only prices the machine type change.
Letting the pool shrink
-
Check how pods are spread and what each node’s requests add up to.
-
Turn on autoscaling for the pool:
Terminal window gcloud container clusters update CLUSTER_NAME --enable-autoscaling \--node-pool=POOL_NAME --min-nodes=MIN_NODES --max-nodes=MAX_NODES --location=LOCATION -
Lower the minimum node count so the autoscaler is allowed to scale down.
-
Consider node auto-provisioning so GKE creates right-sized pools for pending pods.