Skip to main content
schedule · azure

VM scale sets averaging under 5% CPU that should be scaled down on a schedule

resource types
1
rule IDs covered
1
severity
medium

What does ZopNight detect here?

ZopNight flags an Azure virtual machine scale set whose `Percentage CPU` averages below 5% over 30 days, with memory use at or below 60%, low network traffic and no CPU peak near saturation. It recommends scaling the set down during its measured idle window and prices the saving from those idle hours.

Signal and threshold

How ZopNight evaluates VM scale sets averaging under 5% CPU that should be scaled down on a schedule.
Field Value
Rule IDsRC-270
Categoryschedule
Severitymedium
MetricPercentage CPU, Available Memory Percentage, Network In/Out Total
Thresholdavg CPU < 5%, memory used <= 60%, network < 1 MB per minute each way
Evaluation window30d
SourceZopNight
Permissions usedMicrosoft.Compute/virtualMachineScaleSets/read · Microsoft.Insights/Metrics/Read

Capacity that runs all week for a few busy hours

A scale set bills for every instance it keeps running, and its instance count only moves when a rule or a person moves it. Plenty of sets are sized for a peak and then left at that capacity. When the whole set averages under 5% CPU for a month, most of those instance hours are doing nothing.

The fix is not deleting the set. Azure autoscale supports recurring profiles that apply different instance limits on selected days and times, for example scaling in over the weekend, so capacity can follow the calendar.

Measuring a scale set’s utilization

Terminal window
az vmss list --query "[].{name:name, rg:resourceGroup, sku:sku.name, capacity:sku.capacity}" -o table
az monitor metrics list --resource <vmss-resource-id> \
--metric "Percentage CPU" "Available Memory Percentage" "Network In Total" "Network Out Total" \
--aggregation Average Maximum --interval PT1H --offset 30d

Check the maximum column as well as the average. A set that is calm on average but spikes hard once a day needs its capacity at that moment.

Idle thresholds for the whole set

  1. Average CPU over 30 days is below 5%. CPU data is mandatory.
  2. Memory in use averages 60% or less, and average inbound and outbound traffic each stay under 1 MB on Azure Monitor’s per-minute Network In Total and Network Out Total counters.
  3. The CPU peak over the full window does not reach the safe ceiling, so a set whose calm average hides a recurring spike is left alone.
  4. ZopNight has measured idle hours for the set in its activity heatmap, and the set has a monthly cost.

The same 5% floor applies to every scale set, including AKS node pools; names are not used to relax it.

Sets that keep their capacity

A set that already follows a known active schedule is not flagged, and neither is one ZopNight knows has been off for most of the window. Missing pricing or missing idle measurements mean no finding; there is no fixed-percentage estimate. Sets whose problem is their autoscale thresholds rather than their idle hours are covered by Scaling Target Too Low.

Measured idle hours as the saving

Terminal window
monthly saving = scale set monthly cost x measured idle fraction

The recommendation carries the heatmap and a suggested schedule, so the dollar figure and the proposed window describe the same hours.

Scaling the set down on a schedule

  1. Review the suggested start, stop and time zone in the recommendation.
  2. Apply it from the recommendation, or add a recurring autoscale profile with a lower instance count, for example az monitor autoscale profile create --resource-group my-rg --autoscale-name my-autoscale --name weekend --copy-rules default --min-count 1 --count 1 --max-count 1 --recurrence week sat sun --timezone "Pacific Standard Time".
  3. For a one-off reduction, use az vmss scale --resource-group my-rg --name my-vmss --new-capacity 1.
  4. Watch the set through the first scheduled window to confirm it scales back up in time.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·