AKS node pools averaging under 10% CPU, priced for the next smaller VM size
What does ZopNight detect here?
AKS node pools that average under 10% node CPU are carrying far more compute than their pods use. ZopNight flags a pool when `node_cpu_usage_percentage` averages below 10% over the last 30 days, with at least 7 days of data and memory under 50%, and prices the next smaller same-family VM size from real rates.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1357 |
| Category | rightsizing |
| Severity | medium |
| Metric | node_cpu_usage_percentage |
| Threshold | average below 10%, memory below 50% |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | Microsoft.ContainerService/managedClusters/agentPools/read · Microsoft.Insights/Metrics/Read |
Where it applies
A mostly idle pool still bills every node
AKS nodes are ordinary VMs. A pool that averages single-digit CPU across a month is paying for capacity its pods never ask for: too many nodes, nodes that are too large, or both. Pools created for a team that later moved its workloads elsewhere are a common source, as are pools where the node count was raised for an incident and never lowered.
Reading average node CPU for one pool
az monitor metrics list \ --resource /subscriptions/<sub>/resourceGroups/<rg>/providers/Microsoft.ContainerService/managedClusters/<cluster> \ --metric node_cpu_usage_percentage node_memory_working_set_percentage node_disk_usage_percentage \ --aggregation Average --interval PT1H --offset 30d --filter "nodepool eq '<pool>'"
az aks nodepool show --resource-group <rg> --cluster-name <cluster> --name <pool> \ --query "{size:vmSize, count:count, autoscale:enableAutoScaling}"What makes a pool underutilised
- Node CPU data exists with at least 7 days of real coverage, so a freshly created pool is not judged on a few quiet hours.
- Average node CPU over the last 30 days is below 10%.
- If memory data exists, average memory use is below 50%. That check reads the whole series, so a pool that was memory-hot earlier is never scaled down on a calm recent stretch.
- The pool has a price, and real rates exist for its current VM size and for the next smaller size in the same family.
Disk usage is shown on the finding as context but does not affect the decision.
No rate, no recommendation
If either rate is missing there is no finding. The rule never falls back to a flat percentage of the pool’s cost. Pods-per-node counts are not used. Because the test is an average, a pool with a steady low baseline and a short daily spike can still qualify; the peak-based check in AKS Agent Pool VM-Size Rightsizing is the more conservative view, and a pool meeting both gets the same target size and figure from each.
The one-size-down saving
saving = pool monthly cost x (current size rate - smaller size rate) / current size ratecapped at the pool's current costBringing the pool back in line
- Review which pods are scheduled on the pool and whether they could share nodes with another pool.
- Let the cluster autoscaler remove spare nodes. Microsoft’s
autoscaler overview describes it
regularly checking for underutilised nodes and removing them when workloads can be rescheduled:
az aks nodepool update --resource-group <rg> --cluster-name <cluster> --name <pool> --enable-cluster-autoscaler --min-count 1 --max-count <n>. - Or lower the node count directly with
az aks nodepool scale --node-count <n>. - If each node is still mostly idle, move the pool to the next smaller VM size.