AKS node pools running at a fixed size because the cluster autoscaler is off
What does ZopNight detect here?
ZopNight flags an AKS node pool when Azure reports `enableAutoScaling` as false, so the cluster autoscaler never changes its node count. A fixed-size pool pays for its full size through quiet hours and cannot add nodes when pods are stuck pending. ZopNight files it as compliance with no priced saving, since the right bounds depend on the workload.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1356 |
| Category | compliance |
| Severity | medium |
| Metric | none — pure configuration read |
| Threshold | enableAutoScaling = false |
| Source | ZopNight |
| Permissions used | Microsoft.ContainerService/managedClusters/agentPools/read |
Where it applies
A node pool that never resizes
Each AKS node pool is backed by a set of VMs billed while they run. The cluster autoscaler watches for pods that cannot be scheduled because of resource constraints and adds nodes, then removes nodes that stay underused. With it off, the pool keeps whatever count you last set.
That fails in both directions. At night and at weekends the pool bills for capacity no pod
requests. At peak, new pods sit in Pending until someone scales the pool by hand.
Listing pools with autoscaling off
az aks nodepool list --resource-group my-rg --cluster-name my-aks \ --query "[].{pool:name, mode:mode, autoscale:enableAutoScaling, nodes:count}" -o tableRows with autoscale false are fixed-size pools. The mode column separates system pools from
user pools, which usually deserve different bounds.
What ZopNight checks on each pool
The rule evaluates node pools individually, not whole clusters. It reads the pool’s live autoscaling state from Azure and fires only when that state is explicitly false. No utilization metric or window is required, because the finding is about the missing control, not about how busy the pool happens to be today.
Pools the rule does not flag
A pool whose autoscaling state Azure did not report is treated as unknown and skipped. Tags saying the autoscaler is on or off are ignored; only the pool’s real configuration counts. A pool with autoscaling on but a minimum equal to its maximum is not flagged here either, since the autoscaler is technically enabled.
No saving is claimed, only the missing control
This is a compliance finding without a dollar estimate: the saving from autoscaling depends on how far demand falls below today’s fixed count, and on the minimum you choose. The exposure is paying for idle nodes off-peak and failing to schedule pods on-peak.
Enabling the autoscaler on a pool
- Pick bounds. The minimum should cover your baseline and any pod disruption budgets; the maximum caps spend and must fit subnet IP space and vCPU quota.
- Enable it on the existing pool:
az aks nodepool update --resource-group my-rg --cluster-name my-aks \ --name userpool1 --enable-cluster-autoscaler --min-count 1 --max-count 5- Check that pods set realistic CPU and memory requests. Microsoft notes the autoscaler scales up on pending pods rather than on node CPU or memory pressure, so scaling follows what pods request, not what they use.
- Watch scale events for a week and adjust the bounds with
az aks nodepool update --update-cluster-autoscaler.