Skip to main content
rightsizing · azure

AKS node pools whose busiest hour stayed under 40% CPU, fit for the next smaller VM size

resource types
1
rule IDs covered
1
severity
low

What does ZopNight detect here?

AKS node pools bill for every node's VM size, so a pool sized for a peak that never comes pays for idle headroom. ZopNight flags a pool whose `node_cpu_usage_percentage` peak stays under 40% across 30 days, with memory under 50%, and prices a move to the next smaller VM size in the same family from real rates.

Signal and threshold

How ZopNight evaluates AKS node pools whose busiest hour stayed under 40% CPU, fit for the next smaller VM size.
Field Value
Rule IDsRC-1229
Categoryrightsizing
Severitylow
Metricnode_cpu_usage_percentage
Threshold30-day peak below 40%, memory below 50%
Evaluation window30d
SourceZopNight
Permissions usedMicrosoft.ContainerService/managedClusters/agentPools/read · Microsoft.Insights/Metrics/Read

Oversized nodes are paid for around the clock

Each node in an AKS pool is a VM of the pool’s size, billed whether its pods use the capacity or not. Pools are often sized once, for a launch or a load test, and never revisited. If even the busiest hour of the month leaves more than half of every node’s CPU unused, the same workload would fit on the next size down in the same family.

Checking a pool’s peak utilisation

List the pools and their sizes, then read the cluster’s node CPU and memory split by pool. The AKS platform metrics carry a nodepool dimension:

Terminal window
az aks nodepool list --resource-group <rg> --cluster-name <cluster> \
--query "[].{pool:name, size:vmSize, count:count, mode:mode}" -o table
az monitor metrics list \
--resource /subscriptions/<sub>/resourceGroups/<rg>/providers/Microsoft.ContainerService/managedClusters/<cluster> \
--metric node_cpu_usage_percentage node_memory_working_set_percentage \
--aggregation Maximum Average --interval PT1H --offset 30d --filter "nodepool eq '<pool>'"

Gates for a VM-size downsize

  1. The pool’s VM size is known and has a next smaller size in the same family.
  2. That smaller size is allowed for AKS node pools; ZopNight keeps a list of sizes AKS rejects and never proposes them.
  3. Node CPU data covers the full 30 days, and the highest reading in that window is below 40%.
  4. If memory data is present, average memory use is below 50%, so a memory-bound pool is not shrunk.
  5. Real hourly rates exist for both sizes, the smaller one is cheaper, and the pool has a price.

When the rule stays quiet

Less than 30 days of CPU data, no smaller rung, a missing rate or a target size AKS does not allow all mean no finding. Pools named or tagged as production are still evaluated, but the finding is raised at medium severity and asks you to confirm PodDisruptionBudgets and use a maintenance window. A pool that also meets the average-based test in AKS Node Pool Underutilized gets the same target size and figure there; treat the two as one decision.

Pricing the smaller size

Terminal window
saving = pool monthly cost x (current size rate - smaller size rate) / current size rate
cost after fix = pool monthly cost - saving

Resizing the pool without dropping pods

  1. Check that every workload on the pool has a PodDisruptionBudget that allows at least one replica to be evicted.
  2. Resize in place (a preview feature for scale set pools that needs the aks-preview CLI extension): az aks nodepool update --resource-group <rg> --cluster-name <cluster> --name <pool> --node-vm-size <smaller-size>. AKS surges new nodes, then cordons and drains the old ones.
  3. Or follow Microsoft’s recommended method: add a new pool with the smaller size, cordon and drain the old nodes, then delete the old pool.
  4. Watch for pods stuck in Pending after the move, which means the new size is too small.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·