Skip to main content
rightsizing · azure

AKS node pools averaging under 10% CPU, priced for the next smaller VM size

resource types
1
rule IDs covered
1
severity
medium

What does ZopNight detect here?

AKS node pools that average under 10% node CPU are carrying far more compute than their pods use. ZopNight flags a pool when `node_cpu_usage_percentage` averages below 10% over the last 30 days, with at least 7 days of data and memory under 50%, and prices the next smaller same-family VM size from real rates.

Signal and threshold

How ZopNight evaluates AKS node pools averaging under 10% CPU, priced for the next smaller VM size.
Field Value
Rule IDsRC-1357
Categoryrightsizing
Severitymedium
Metricnode_cpu_usage_percentage
Thresholdaverage below 10%, memory below 50%
Evaluation window30d
SourceZopNight
Permissions usedMicrosoft.ContainerService/managedClusters/agentPools/read · Microsoft.Insights/Metrics/Read

A mostly idle pool still bills every node

AKS nodes are ordinary VMs. A pool that averages single-digit CPU across a month is paying for capacity its pods never ask for: too many nodes, nodes that are too large, or both. Pools created for a team that later moved its workloads elsewhere are a common source, as are pools where the node count was raised for an incident and never lowered.

Reading average node CPU for one pool

Terminal window
az monitor metrics list \
--resource /subscriptions/<sub>/resourceGroups/<rg>/providers/Microsoft.ContainerService/managedClusters/<cluster> \
--metric node_cpu_usage_percentage node_memory_working_set_percentage node_disk_usage_percentage \
--aggregation Average --interval PT1H --offset 30d --filter "nodepool eq '<pool>'"
az aks nodepool show --resource-group <rg> --cluster-name <cluster> --name <pool> \
--query "{size:vmSize, count:count, autoscale:enableAutoScaling}"

What makes a pool underutilised

  • Node CPU data exists with at least 7 days of real coverage, so a freshly created pool is not judged on a few quiet hours.
  • Average node CPU over the last 30 days is below 10%.
  • If memory data exists, average memory use is below 50%. That check reads the whole series, so a pool that was memory-hot earlier is never scaled down on a calm recent stretch.
  • The pool has a price, and real rates exist for its current VM size and for the next smaller size in the same family.

Disk usage is shown on the finding as context but does not affect the decision.

No rate, no recommendation

If either rate is missing there is no finding. The rule never falls back to a flat percentage of the pool’s cost. Pods-per-node counts are not used. Because the test is an average, a pool with a steady low baseline and a short daily spike can still qualify; the peak-based check in AKS Agent Pool VM-Size Rightsizing is the more conservative view, and a pool meeting both gets the same target size and figure from each.

The one-size-down saving

Terminal window
saving = pool monthly cost x (current size rate - smaller size rate) / current size rate
capped at the pool's current cost

Bringing the pool back in line

  1. Review which pods are scheduled on the pool and whether they could share nodes with another pool.
  2. Let the cluster autoscaler remove spare nodes. Microsoft’s autoscaler overview describes it regularly checking for underutilised nodes and removing them when workloads can be rescheduled: az aks nodepool update --resource-group <rg> --cluster-name <cluster> --name <pool> --enable-cluster-autoscaler --min-count 1 --max-count <n>.
  3. Or lower the node count directly with az aks nodepool scale --node-count <n>.
  4. If each node is still mostly idle, move the pool to the next smaller VM size.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·