Azure ML compute clusters holding warm nodes that ran no jobs for 30 days
What does ZopNight detect here?
Azure Machine Learning compute clusters with a minimum node count above 0 keep those nodes allocated and billed between jobs. ZopNight flags a cluster when the `Active Nodes` metric shows no node running a job for 30 days while `Total Nodes` proves the warm floor was allocated, and prices the idle minimum nodes.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-223 |
| Category | idle |
| Severity | medium |
| Metric | Active Nodes |
| Threshold | min nodes above 0 and zero active nodes |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | Microsoft.MachineLearningServices/workspaces/computes/read · Microsoft.Insights/Metrics/Read |
Where it applies
Where a warm training cluster leaks money
An AmlCompute cluster scales between a minimum and a maximum node count. Microsoft’s compute cluster guide is direct about the cost: to avoid charges when no jobs are running, set the minimum nodes to 0, because any value larger than 0 keeps that number of nodes running even when they are not in use.
Teams often raise the minimum during a sprint of experiments to skip the allocation wait, then forget it. The GPU or CPU nodes stay allocated, and billed, long after the last job.
Spotting clusters with a non-zero minimum
List the clusters in a workspace and look for a minimum above zero:
az ml compute list --type AmlCompute \ --resource-group <rg> --workspace-name <workspace> \ --query '[?min_instances > `0`].{name:name, min:min_instances, size:size}' -o tableThen confirm whether any node actually ran work. The workspace exposes Active Nodes (nodes
running a job) and Total Nodes in Azure Monitor:
az monitor metrics list \ --resource /subscriptions/<sub>/resourceGroups/<rg>/providers/Microsoft.MachineLearningServices/workspaces/<workspace> \ --metric "Active Nodes" "Total Nodes" --aggregation Maximum --interval PT1H --offset 30dProof required before a cluster is called idle
- The cluster’s configured minimum node count is greater than 0.
- A full 30 days of
Active Nodesdata exists, and both its average and its peak stay below one node: no job ran on any node in that time. - A full 30 days of
Total Nodesdata exists and peaks at one node or more, which proves the warm floor was really allocated and billing. - The cluster has a positive monthly price.
Signals that keep a cluster off the list
If either node series is missing or covers less than 30 days, there is no finding. A cluster
whose minimum is already 0 is never flagged, since it scales to nothing on its own. CpuUtilization
is only emitted while a job runs, so its absence proves nothing; when it is present and shows an
average or peak of 5% or more, that busy signal vetoes the finding even if the node counts look idle.
Managed online endpoints are a different resource and are covered by Azure ML Online Deployment Idle.
Pricing only the warm nodes
The saving is the slice of the cluster’s cost that belongs to its minimum nodes, never more than the cluster costs today:
per-node cost = cluster monthly cost / current node count (at least 1)saving = min(per-node cost x minimum nodes, cluster monthly cost)Letting the cluster scale to zero
- Check Azure Machine Learning studio for running or queued jobs on the cluster.
- Set the minimum to zero, either in studio (Compute, Compute clusters, Edit) or with
az ml compute update --name <cluster> --min-instances 0 --resource-group <rg> --workspace-name <workspace>. - Keep a sensible idle time before scale-down so back-to-back jobs can reuse nodes.