Skip to main content
resource · azure

Azure ML Compute Cluster

live rule families
1
schedulable
yes
category
ai-ml-services

Does ZopNight manage Azure ML Compute Cluster?

Azure ML compute clusters bill for every node they hold, and a minimum node count above 0 keeps that many VMs allocated even with no jobs queued. ZopNight schedules clusters by setting the minimum to 0 outside training windows (idle nodes deallocate) and restores the saved minimum on start.

Rules that fire on Azure ML Compute Cluster

At a glance

Azure ML Compute Cluster coverage facts.
Field Value
Scheduling notesminimum node count is set to 0 on stop (idle nodes deallocate) and the saved minimum is restored on start.

ML compute clusters are autoscaling VM pools for training jobs. Clusters configured with a minimum node count above zero hold those nodes, and their bill, even when no jobs run.

The minimum node count is the meter that matters

A compute cluster bills per node-hour for the VMs it currently holds. The autoscaler already releases nodes above the minimum when the queue empties, so the standing cost of a cluster is set entirely by its minimum node count. A minimum of zero means an empty queue costs nothing; a minimum of two on a GPU size means two expensive VMs bill through every night and weekend, warm for jobs that are not coming.

Cluster signals in ZopNight’s inventory

Discovered via the AML enricher with min/max node configuration. Cost Management billing attributes spend, recommendations flag nonzero minimum node counts, and ML auto-tagging covers this type. Job discovery links each run to its target cluster, so a cluster’s configured floor can be judged against what actually executes on it.

What a scheduled scale-down actually changes

On stop, ZopNight sets the minimum node count to 0 so idle nodes deallocate, and it saves the configured minimum; on start, the saved value is restored. Running jobs are untouched. Nodes executing a job are kept until the job completes, so the change affects idle capacity only. A training run that spills past the schedule boundary finishes normally; what disappears is the warm floor.

How training clusters overspend

Two habits dominate. Teams raise the minimum during a crunch to skip node spin-up delays, then never lower it. And clusters get sized on premium GPU SKUs for one experiment, after which the same floor idles at GPU rates. Both leaks are invisible in the job history and obvious in the node configuration.

Verifying a cluster’s floor in ML studio

Azure ML studio → Compute → Compute clusters shows each cluster’s minimum and maximum nodes and current allocation; the minimum column is the standing bill.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

417 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

417 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·