# AKS Agent Pool VM-Size Rightsizing

> Flags AKS node pools peaking under 40% CPU for 30 days and prices the next smaller same-family VM size from real rates.

Source: https://zop.dev/integrations/azure/recommendations/aks-agent-pool-vm-size-rightsizing

---

## Oversized nodes are paid for around the clock

Each node in an AKS pool is a VM of the pool's size, billed whether its pods use the capacity or
not. Pools are often sized once, for a launch or a load test, and never revisited. If even the
busiest hour of the month leaves more than half of every node's CPU unused, the same workload would
fit on the next size down in the same family.

## Checking a pool's peak utilisation

List the pools and their sizes, then read the cluster's node CPU and memory split by pool. The AKS
platform metrics carry a `nodepool` dimension:

```bash
az aks nodepool list --resource-group <rg> --cluster-name <cluster> \
  --query "[].{pool:name, size:vmSize, count:count, mode:mode}" -o table

az monitor metrics list \
  --resource /subscriptions/<sub>/resourceGroups/<rg>/providers/Microsoft.ContainerService/managedClusters/<cluster> \
  --metric node_cpu_usage_percentage node_memory_working_set_percentage \
  --aggregation Maximum Average --interval PT1H --offset 30d --filter "nodepool eq '<pool>'"
```

## Gates for a VM-size downsize

1. The pool's VM size is known and has a next smaller size in the same family.
2. That smaller size is allowed for AKS node pools; ZopNight keeps a list of sizes AKS rejects and
   never proposes them.
3. Node CPU data covers the full 30 days, and the highest reading in that window is below 40%.
4. If memory data is present, average memory use is below 50%, so a memory-bound pool is not shrunk.
5. Real hourly rates exist for both sizes, the smaller one is cheaper, and the pool has a price.

## When the rule stays quiet

Less than 30 days of CPU data, no smaller rung, a missing rate or a target size AKS does not allow
all mean no finding. Pools named or tagged as production are still evaluated, but the finding is
raised at medium severity and asks you to confirm PodDisruptionBudgets and use a maintenance window.
A pool that also meets the average-based test in
<a href="https://zop.dev/integrations/azure/recommendations/aks-node-pool-underutilized">AKS Node Pool Underutilized</a>
gets the same target size and figure there; treat the two as one decision.

## Pricing the smaller size

```text
saving = pool monthly cost x (current size rate - smaller size rate) / current size rate
cost after fix = pool monthly cost - saving
```

## Resizing the pool without dropping pods

1. Check that every workload on the pool has a PodDisruptionBudget that allows at least one replica
   to be evicted.
2. Resize in place (a preview feature for scale set pools that needs the `aks-preview` CLI extension):
   `az aks nodepool update --resource-group <rg> --cluster-name <cluster> --name <pool> --node-vm-size <smaller-size>`.
   AKS surges new nodes, then cordons and drains the old ones.
3. Or follow Microsoft's
   [recommended method](https://learn.microsoft.com/en-us/azure/aks/resize-node-pool): add a new
   pool with the smaller size, cordon and drain the old nodes, then delete the old pool.
4. Watch for pods stuck in `Pending` after the move, which means the new size is too small.

**Note**
System node pools have extra limits: Microsoft says sizes with fewer than two vCPUs and 4 GB of memory may not work, and B-series VMs are not supported for system pools.
