Azure Machine Learning online deployments that served zero requests
What does ZopNight detect here?
ZopNight flags an Azure Machine Learning managed online deployment whose `RequestsPerMinute` metric shows zero on average and at peak, with at least 7 days of data and a known cost. A managed online deployment keeps its instances, often GPU VMs, running for as long as it exists, so a model endpoint nobody calls bills the full instance rate.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1397 |
| Category | idle |
| Severity | high |
| Metric | RequestsPerMinute |
| Threshold | average and maximum = 0 |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | Microsoft.MachineLearningServices/workspaces/onlineEndpoints/read · Microsoft.MachineLearningServices/workspaces/onlineEndpoints/deployments/read · Microsoft.Insights/Metrics/Read |
Where it applies
Instances bill for as long as the deployment exists
A managed online endpoint is the address; each deployment behind it is a set of instances of a chosen VM size running your model. You pick the instance count, and Azure Machine Learning also reserves 20% extra quota on some VM sizes for upgrades, per the online endpoints overview. Those instances run whether or not requests arrive.
Deployments are also easy to lose track of in the bill. Microsoft’s
cost guide for online endpoints
explains that costs accrue to the workspace and must be filtered by the azuremlendpoint and
azuremldeployment tags to see a single deployment.
Checking requests on a deployment
az ml online-deployment list --endpoint-name my-endpoint \ --resource-group my-rg --workspace-name my-workspace -o table
az monitor metrics list --resource <deployment-resource-id> \ --metric RequestsPerMinute CpuUtilizationPercentage GpuUtilizationPercentage \ --offset 30d --interval PT24H --aggregation Average MaximumThe az ml commands come from the Azure CLI ml extension.
Evidence needed for a delete recommendation
- The
RequestsPerMinuteseries exists for the deployment. - Its average is 0 and its maximum is 0.
- It covers at least 7 days, so a deployment created this week is not flagged before callers are pointed at it.
- The deployment has a known cost above zero. Billing for a new deployment can take a day or more to appear, so the finding may show up on a later pass rather than immediately.
CPU and GPU utilization are attached as supporting evidence when available. They do not decide the outcome.
Deployments that are not flagged
No request data means no finding: absence of the metric is not treated as zero traffic. A single request in the window clears the deployment. Deployments with no cost yet recorded are also skipped, because a finding without a price cannot be ranked against anything else.
Saving is the deployment’s full run rate
saving = current monthly cost of the deployment's instancescost after fix = 0Removing an unused deployment
- Confirm no application, pipeline or traffic rule routes to the deployment; check the endpoint’s traffic split in Azure Machine Learning studio under Endpoints.
- Delete it:
az ml online-deployment delete --name my-deployment --endpoint-name my-endpoint --resource-group my-rg --workspace-name my-workspace --yes. - If the endpoint has no deployments left, delete it too with
az ml online-endpoint delete. - For models that need occasional inference, consider autoscale on the deployment through Azure Monitor autoscale, so instance count follows demand.