Skip to main content
schedule · azure

Non-production Azure VMs with regular idle hours that a start and stop schedule could recover

resource types
1
rule IDs covered
1
severity
medium

What does ZopNight detect here?

ZopNight flags a non-production Azure VM when its hourly `Percentage CPU` and network pattern over the last 21 days shows at least 20% of the week could be switched off by a single start and stop schedule. The saving is the VM's cost multiplied by the hours the suggested schedule would actually turn it off.

Signal and threshold

How ZopNight evaluates Non-production Azure VMs with regular idle hours that a start and stop schedule could recover.
Field Value
Rule IDsRC-213
Categoryschedule
Severitymedium
MetricPercentage CPU, Network In Total, Network Out Total
Thresholdrecoverable hours >= 20% and idle < 95%
Evaluation window21d
SourceZopNight
Permissions usedMicrosoft.Compute/virtualMachines/read · Microsoft.Resources/subscriptions/resourceGroups/read · Microsoft.Insights/Metrics/Read

Running dev and test VMs through nights and weekends

A running VM is billed by the hour and a deallocated one is not. Microsoft’s VM states and billing page marks the deallocated state as not billed for instance usage, while a VM stopped from inside the guest stays allocated and keeps billing. A development VM used from nine to six, Monday to Friday, is busy for about a quarter of the week and pays for all of it.

Looking at a VM’s weekly pattern yourself

Terminal window
az vm list -d \
--query "[].{name:name, rg:resourceGroup, power:powerState, env:tags.environment}" \
-o table
az monitor metrics list --resource <vm-resource-id> \
--metric "Percentage CPU" "Network In Total" "Network Out Total" \
--aggregation Average --interval PT1H --offset 21d

Lay the hourly values out by day of week and hour. Blocks of consistently quiet hours are what a schedule can recover.

Building the heatmap and the schedule

  1. The VM is provisioned successfully and is positively non-production: a dev or test environment tag, name or resource group, with a production tag as an absolute veto. A neutral name such as app-server-07 is not assumed to be non-production.
  2. VMs with critical or stateful roles, spot or low-priority VMs, scale set members, Databricks VMs, and VMs covered by a reservation or savings plan are skipped; stopping a reserved VM does not reduce the commitment.
  3. The data has to be current and complete enough: the last datapoint no more than 3 days old, at least 18 days of span with 7 or more observed days, and no gap longer than 3 days. Only the last 21 days feed the heatmap.
  4. An hour counts as active when CPU is at least 10%, or when combined network traffic is above a quiet line set by the VM’s own busy periods (a quarter of its 90th-percentile traffic, with a fixed floor for background agent noise).
  5. A start and stop schedule is derived from the active hours, with one hour of lead time before the first active hour each day.

When no schedule is proposed

The rule fires only when the hours the schedule would switch off are at least 20% of the week. A VM that is idle 95% of the time or more is not a scheduling case at all; it belongs with Idle Azure VM. Stale metrics, too little history, missing pricing, or a pattern with no clear active window all mean no finding.

Saving based on hours the schedule turns off

Terminal window
monthly saving = VM monthly cost x (running hours outside the suggested window / 168)

The price describes exactly what the suggested schedule does, not the raw idle share, so hours of idleness scattered through the working day do not inflate it.

Applying the start and stop schedule

  1. Review the suggested start time, stop time and time zone in the recommendation.
  2. Check with the VM’s owners for out-of-hours jobs the heatmap might have read as quiet.
  3. Apply the schedule from the recommendation, and confirm the VM shows as deallocated, not just stopped, outside its window.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·