Azure ML Job
Does ZopNight manage Azure ML Job?
Azure ML jobs are transient training, pipeline, or scoring runs that bill through the compute they occupy, not as standing resources. ZopNight records each job's status and run duration via the AML enricher, grouped under a Jobs node named after its workspace, so runs stuck in a running state are easy to spot.
Rules that fire on Azure ML Job
No active rule family targets Azure ML Job today. Rules that used to are retired, and retired rules publish no pages and fire no findings. Scheduling and permissions coverage are unaffected.
At a glance
| Field | Value |
|---|---|
| Scheduling notes | discovery only; jobs are transient executions rather than standing infrastructure. |
ML jobs are individual training, pipeline, or scoring runs executed on ML compute. Jobs reveal which clusters are actually earning their keep and which are idle allocations.
A run, not a resource: how job cost accrues
A job has no meter of its own. Its cost is the node-hours it occupies on the compute it targets: a compute cluster’s VMs while the run executes, released when it finishes. That indirection is why job spend never appears as a “job” line anywhere: it lands on the cluster. The corollary is that a runaway job is a compute leak wearing a workload’s name, billing cluster nodes for as long as it stays in a running state.
What ZopNight records for each run
Discovered via the AML enricher with status and run timing (start, end and duration), grouped under a per-workspace Jobs node. The record does not capture which cluster hosted the run, and job history does not feed ML compute recommendations. To catch a cluster whose job history is empty for weeks while it holds a warm minimum, compare the target compute column in ML studio (below) with the cluster list.
Why jobs are discovery-only
Jobs are transient executions rather than standing infrastructure, so there is nothing for a schedule to stop and restart: pausing a half-finished training run is not a savings action, it is a broken experiment. ZopNight therefore treats jobs as read-only telemetry, and applies its scheduling to the clusters and instances the jobs run on.
Job-shaped waste to watch for
The patterns worth hunting: runs stuck in a running state long after their expected duration, holding nodes; retry loops resubmitting a failing pipeline against an expensive cluster; and experiment sweeps launched at a parallelism the team never reviewed. Each shows up first in the job list, not the cost report.
Reading the run history in ML studio
Azure ML studio → Jobs lists runs with status, duration, and target compute; sort by duration to find the runs that held nodes longest.