Skip to main content
idle · azure

Azure ML compute clusters holding warm nodes that ran no jobs for 30 days

resource types
1
rule IDs covered
1
severity
medium

What does ZopNight detect here?

Azure Machine Learning compute clusters with a minimum node count above 0 keep those nodes allocated and billed between jobs. ZopNight flags a cluster when the `Active Nodes` metric shows no node running a job for 30 days while `Total Nodes` proves the warm floor was allocated, and prices the idle minimum nodes.

Signal and threshold

How ZopNight evaluates Azure ML compute clusters holding warm nodes that ran no jobs for 30 days.
Field Value
Rule IDsRC-223
Categoryidle
Severitymedium
MetricActive Nodes
Thresholdmin nodes above 0 and zero active nodes
Evaluation window30d
SourceZopNight
Permissions usedMicrosoft.MachineLearningServices/workspaces/computes/read · Microsoft.Insights/Metrics/Read

Where a warm training cluster leaks money

An AmlCompute cluster scales between a minimum and a maximum node count. Microsoft’s compute cluster guide is direct about the cost: to avoid charges when no jobs are running, set the minimum nodes to 0, because any value larger than 0 keeps that number of nodes running even when they are not in use.

Teams often raise the minimum during a sprint of experiments to skip the allocation wait, then forget it. The GPU or CPU nodes stay allocated, and billed, long after the last job.

Spotting clusters with a non-zero minimum

List the clusters in a workspace and look for a minimum above zero:

Terminal window
az ml compute list --type AmlCompute \
--resource-group <rg> --workspace-name <workspace> \
--query '[?min_instances > `0`].{name:name, min:min_instances, size:size}' -o table

Then confirm whether any node actually ran work. The workspace exposes Active Nodes (nodes running a job) and Total Nodes in Azure Monitor:

Terminal window
az monitor metrics list \
--resource /subscriptions/<sub>/resourceGroups/<rg>/providers/Microsoft.MachineLearningServices/workspaces/<workspace> \
--metric "Active Nodes" "Total Nodes" --aggregation Maximum --interval PT1H --offset 30d

Proof required before a cluster is called idle

  • The cluster’s configured minimum node count is greater than 0.
  • A full 30 days of Active Nodes data exists, and both its average and its peak stay below one node: no job ran on any node in that time.
  • A full 30 days of Total Nodes data exists and peaks at one node or more, which proves the warm floor was really allocated and billing.
  • The cluster has a positive monthly price.

Signals that keep a cluster off the list

If either node series is missing or covers less than 30 days, there is no finding. A cluster whose minimum is already 0 is never flagged, since it scales to nothing on its own. CpuUtilization is only emitted while a job runs, so its absence proves nothing; when it is present and shows an average or peak of 5% or more, that busy signal vetoes the finding even if the node counts look idle.

Managed online endpoints are a different resource and are covered by Azure ML Online Deployment Idle.

Pricing only the warm nodes

The saving is the slice of the cluster’s cost that belongs to its minimum nodes, never more than the cluster costs today:

Terminal window
per-node cost = cluster monthly cost / current node count (at least 1)
saving = min(per-node cost x minimum nodes, cluster monthly cost)

Letting the cluster scale to zero

  1. Check Azure Machine Learning studio for running or queued jobs on the cluster.
  2. Set the minimum to zero, either in studio (Compute, Compute clusters, Edit) or with az ml compute update --name <cluster> --min-instances 0 --resource-group <rg> --workspace-name <workspace>.
  3. Keep a sensible idle time before scale-down so back-to-back jobs can reuse nodes.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·