AKS node pools whose busiest hour stayed under 40% CPU, fit for the next smaller VM size
What does ZopNight detect here?
AKS node pools bill for every node's VM size, so a pool sized for a peak that never comes pays for idle headroom. ZopNight flags a pool whose `node_cpu_usage_percentage` peak stays under 40% across 30 days, with memory under 50%, and prices a move to the next smaller VM size in the same family from real rates.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1229 |
| Category | rightsizing |
| Severity | low |
| Metric | node_cpu_usage_percentage |
| Threshold | 30-day peak below 40%, memory below 50% |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | Microsoft.ContainerService/managedClusters/agentPools/read · Microsoft.Insights/Metrics/Read |
Where it applies
Oversized nodes are paid for around the clock
Each node in an AKS pool is a VM of the pool’s size, billed whether its pods use the capacity or not. Pools are often sized once, for a launch or a load test, and never revisited. If even the busiest hour of the month leaves more than half of every node’s CPU unused, the same workload would fit on the next size down in the same family.
Checking a pool’s peak utilisation
List the pools and their sizes, then read the cluster’s node CPU and memory split by pool. The AKS
platform metrics carry a nodepool dimension:
az aks nodepool list --resource-group <rg> --cluster-name <cluster> \ --query "[].{pool:name, size:vmSize, count:count, mode:mode}" -o table
az monitor metrics list \ --resource /subscriptions/<sub>/resourceGroups/<rg>/providers/Microsoft.ContainerService/managedClusters/<cluster> \ --metric node_cpu_usage_percentage node_memory_working_set_percentage \ --aggregation Maximum Average --interval PT1H --offset 30d --filter "nodepool eq '<pool>'"Gates for a VM-size downsize
- The pool’s VM size is known and has a next smaller size in the same family.
- That smaller size is allowed for AKS node pools; ZopNight keeps a list of sizes AKS rejects and never proposes them.
- Node CPU data covers the full 30 days, and the highest reading in that window is below 40%.
- If memory data is present, average memory use is below 50%, so a memory-bound pool is not shrunk.
- Real hourly rates exist for both sizes, the smaller one is cheaper, and the pool has a price.
When the rule stays quiet
Less than 30 days of CPU data, no smaller rung, a missing rate or a target size AKS does not allow all mean no finding. Pools named or tagged as production are still evaluated, but the finding is raised at medium severity and asks you to confirm PodDisruptionBudgets and use a maintenance window. A pool that also meets the average-based test in AKS Node Pool Underutilized gets the same target size and figure there; treat the two as one decision.
Pricing the smaller size
saving = pool monthly cost x (current size rate - smaller size rate) / current size ratecost after fix = pool monthly cost - savingResizing the pool without dropping pods
- Check that every workload on the pool has a PodDisruptionBudget that allows at least one replica to be evicted.
- Resize in place (a preview feature for scale set pools that needs the
aks-previewCLI extension):az aks nodepool update --resource-group <rg> --cluster-name <cluster> --name <pool> --node-vm-size <smaller-size>. AKS surges new nodes, then cordons and drains the old ones. - Or follow Microsoft’s recommended method: add a new pool with the smaller size, cordon and drain the old nodes, then delete the old pool.
- Watch for pods stuck in
Pendingafter the move, which means the new size is too small.