VM scale sets averaging under 5% CPU that should be scaled down on a schedule
What does ZopNight detect here?
ZopNight flags an Azure virtual machine scale set whose `Percentage CPU` averages below 5% over 30 days, with memory use at or below 60%, low network traffic and no CPU peak near saturation. It recommends scaling the set down during its measured idle window and prices the saving from those idle hours.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-270 |
| Category | schedule |
| Severity | medium |
| Metric | Percentage CPU, Available Memory Percentage, Network In/Out Total |
| Threshold | avg CPU < 5%, memory used <= 60%, network < 1 MB per minute each way |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | Microsoft.Compute/virtualMachineScaleSets/read · Microsoft.Insights/Metrics/Read |
Where it applies
Capacity that runs all week for a few busy hours
A scale set bills for every instance it keeps running, and its instance count only moves when a rule or a person moves it. Plenty of sets are sized for a peak and then left at that capacity. When the whole set averages under 5% CPU for a month, most of those instance hours are doing nothing.
The fix is not deleting the set. Azure autoscale supports recurring profiles that apply different instance limits on selected days and times, for example scaling in over the weekend, so capacity can follow the calendar.
Measuring a scale set’s utilization
az vmss list --query "[].{name:name, rg:resourceGroup, sku:sku.name, capacity:sku.capacity}" -o table
az monitor metrics list --resource <vmss-resource-id> \ --metric "Percentage CPU" "Available Memory Percentage" "Network In Total" "Network Out Total" \ --aggregation Average Maximum --interval PT1H --offset 30dCheck the maximum column as well as the average. A set that is calm on average but spikes hard once a day needs its capacity at that moment.
Idle thresholds for the whole set
- Average CPU over 30 days is below 5%. CPU data is mandatory.
- Memory in use averages 60% or less, and average inbound and outbound traffic each stay under 1 MB on Azure Monitor’s per-minute Network In Total and Network Out Total counters.
- The CPU peak over the full window does not reach the safe ceiling, so a set whose calm average hides a recurring spike is left alone.
- ZopNight has measured idle hours for the set in its activity heatmap, and the set has a monthly cost.
The same 5% floor applies to every scale set, including AKS node pools; names are not used to relax it.
Sets that keep their capacity
A set that already follows a known active schedule is not flagged, and neither is one ZopNight knows has been off for most of the window. Missing pricing or missing idle measurements mean no finding; there is no fixed-percentage estimate. Sets whose problem is their autoscale thresholds rather than their idle hours are covered by Scaling Target Too Low.
Measured idle hours as the saving
monthly saving = scale set monthly cost x measured idle fractionThe recommendation carries the heatmap and a suggested schedule, so the dollar figure and the proposed window describe the same hours.
Scaling the set down on a schedule
- Review the suggested start, stop and time zone in the recommendation.
- Apply it from the recommendation, or add a recurring autoscale profile with a lower instance
count, for example
az monitor autoscale profile create --resource-group my-rg --autoscale-name my-autoscale --name weekend --copy-rules default --min-count 1 --count 1 --max-count 1 --recurrence week sat sun --timezone "Pacific Standard Time". - For a one-off reduction, use
az vmss scale --resource-group my-rg --name my-vmss --new-capacity 1. - Watch the set through the first scheduled window to confirm it scales back up in time.