VM scale sets whose CPU scale-out rule waits until utilization passes 90%
What does ZopNight detect here?
ZopNight flags an Azure virtual machine scale set when the `Percentage CPU` threshold on its autoscale scale-out rule is set above 90. By the time average CPU crosses that line, the running instances are already close to saturation, and new ones still need time to boot. ZopNight recommends lowering the threshold to 70.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-ASC-004 |
| Category | compliance |
| Severity | high |
| Metric | Percentage CPU (autoscale threshold) |
| Threshold | scale-out threshold > 90 |
| Source | ZopNight |
| Permissions used | Microsoft.Compute/virtualMachineScaleSets/read · Microsoft.Insights/AutoscaleSettings/Read |
Where it applies
Why a 95% scale-out threshold arrives too late
An autoscale rule compares a metric with a threshold over a timeWindow before it acts. In the
autoscale settings schema,
Microsoft’s own example scale-out rule fires when the scale set’s average Percentage CPU is
greater than 85 over the past 10 minutes. That averaging exists to smooth out spikes, but it also
means the rule only reacts after a sustained period at the threshold.
Set the threshold at 95 and the instances have to sit near full CPU for the whole window before autoscale even starts adding capacity. The new VMs then boot and warm up while the existing ones queue requests, time out health probes or throttle. The fleet is not over-provisioned in this state, it is under-protected: the setting saves little money and spends reliability.
Reading the CPU threshold on a scale set’s autoscale setting
Find the autoscale setting that targets the scale set, then print the threshold of every rule that scales out on CPU:
az monitor autoscale list --resource-group my-rg \ --query "[].{name:name, target:targetResourceUri}" -o table
az monitor autoscale show --resource-group my-rg --name my-autoscale \ --query "profiles[].rules[?metricTrigger.metricName=='Percentage CPU' && scaleAction.direction=='Increase'].metricTrigger.threshold" \ -o tsvA value above 90 is what this rule reports. az monitor autoscale rule list --resource-group my-rg --autoscale-name my-autoscale
shows the full rule, including its timeWindow and timeAggregation.
The 90% line and what ZopNight reads
ZopNight takes the Percentage CPU scale-out threshold from the autoscale setting attached to
the scale set and raises a finding when all of these hold:
- The scale set is provisioned successfully or running.
- A numeric threshold was read from the autoscale setting.
- That threshold is strictly greater than 90.
A threshold of exactly 90 passes. The finding reports the value it read and proposes 70 as the new target.
Scale sets this check leaves alone
When the autoscale setting cannot be read, or the threshold is missing or not a number, ZopNight does not guess and does not raise a finding. A scale set with no autoscale setting at all is reported by VMSS Autoscale Setting Not Configured instead. The opposite mistake, a threshold so low that the set scales out while mostly idle, is Scaling Target Too Low, and rules that can re-fire too quickly are covered by Scaling Cooldown Too Short.
A reliability finding with no dollar figure
This is a compliance check and carries no savings estimate. Lowering the threshold can add instance hours, because the set scales out earlier. What you buy with them is time: capacity arrives while the existing instances still have room to absorb load, rather than after users have already felt the slowdown.
Lowering the scale-out threshold to 70%
- Open the scale set’s autoscale setting (portal: the scale set, then Scaling), or export it
with
az monitor autoscale show. - Change the CPU scale-out rule’s threshold to 70, for example by recreating it with
az monitor autoscale rule create --resource-group my-rg --autoscale-name my-autoscale --condition "Percentage CPU > 70 avg 5m" --scale out 1and deleting the old rule. - Keep a matching scale-in rule on the same metric with a clear gap below it. Microsoft’s autoscale best practices warn that mixing metrics across scale-out and scale-in rules can cause flapping.
- Watch scale events in the activity log for a few days and confirm the new instance count is stable.