Skip to main content
resource · azure

Azure ML Online Deployment

live rule families
1
schedulable
no
category
ai-ml-services

Does ZopNight manage Azure ML Online Deployment?

Azure ML online deployments hold a fixed set of CPU or GPU instances that bill every hour of every day, whether 1 request or 1 million arrives. ZopNight captures each deployment's instance type and count, attributes its spend from Cost Management, and flags over-provisioned inference capacity in recommendations.

Rules that fire on Azure ML Online Deployment

At a glance

Azure ML Online Deployment coverage facts.
Field Value
Scheduling notesdiscovery and cost visibility only.

Online deployments are the provisioned instance sets behind real-time inference endpoints, billed per instance-hour continuously. Over-instanced deployments for low-traffic models burn GPU or CPU money around the clock.

Instance-hours are the whole story

A deployment’s bill is its instance count times its instance size, accruing every hour it exists. Traffic volume changes nothing about the meter: a model scoring a handful of requests a day on three GPU instances pays exactly what a saturated one pays. That makes deployment sizing the single decision that sets real-time inference cost, and makes low-traffic production models the place where the gap between provisioned and needed is widest.

What ZopNight knows about each deployment

Discovered via the AML enricher with instance type and count, under the endpoint that routes to it. Cost Management billing attributes spend per deployment, and recommendations flag over-provisioned inference capacity: deployments whose instance count outruns any traffic they see.

Why deployments are visibility-only for now

A deployment cannot be stopped without dropping the model out of its endpoint’s traffic split, so ZopNight scopes it to discovery and cost visibility rather than scheduling. The savings lever is structural: shrink the instance count, move to a smaller SKU, or retire the deployment. All three are decisions the per-deployment cost attribution is designed to force into view.

Where inference capacity goes to waste

Three recurring shapes: deployments provisioned “for launch traffic” that never materialized; the second half of an A/B test still holding full capacity months after the decision; and staging deployments mirroring production instance counts on models nobody calls after sign-off. GPU deployments deserve the first pass in any review, since a single over-provisioned instance there outweighs several CPU ones.

Checking a deployment’s size in ML studio

Azure ML studio → Endpoints → open the endpoint → Deployments tab shows each deployment’s instance type and count next to its share of traffic. Instance counts that dwarf traffic share are the audit finding.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

417 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

417 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·