Machine learning compute instances run up cost around the clock, and in non-production that is mostly waste. ML Compute instances typically bill in the range of $100-5,000/mo per instance, and most of that meter runs on hours no one is working.
ZopNight optimizes Azure ML Compute on two fronts at once: it schedules ML Compute to stop outside working hours, and it surfaces resources that are the wrong size or genuinely idle. Every figure it reports is measured against your actual usage, not a projection, which is the same FinOps discipline that separates it from dashboard-first tools like CloudHealth.
Start with the calendar: schedule the idle hours
The single biggest ML Compute win is refusing to pay for nights and weekends. You define a cron once, stop in the evening, start before the team logs on, and ZopNight applies the Stop/Start to every resource in scope, in dependency order, with overrides for the rare late deploy.
This is the highest-ROI, lowest-risk action available, which is why it comes first. There is no architectural change, no code to modify, and production stays excluded. Most teams see the first savings in the very first cycle.
Then find what’s idle even during the day
Scheduling handles the predictable waste; idle detection handles the surprises. ZopNight analyzes CPU, network, and connection metrics over a rolling window to distinguish a genuinely idle ML Compute instance from one with legitimate but intermittent traffic.
Each finding comes with a dollar estimate, so you act on the biggest first, the forgotten test box, the proof-of-concept that outlived its project, the environment someone spun up and never tore down.
Finally, rightsize what’s oversized
When a resource is used but over-provisioned, scheduling is the wrong tool, rightsizing is. ZopNight’s guided recommendations move an over-sized ML Compute instance to the correct type using measured utilization over a real window, never a flat 24×7 assumption. Because the saving is computed from what the resource actually did, the headline number matches the bill you will actually see.
Optimize ML Compute by use case
ML Compute cost work breaks into a few jobs. Go deeper on any of them:
Cost Optimization, Resource Scheduling, Idle Resource Detection, Rightsizing, FinOps Automation, Orphan Resource Cleanup, Cloud Cost Reporting, Multi-Cloud Management, Event Readiness, Cost Anomaly Detection, Showback and Cost Attribution, Smart Tags, Autoscaler Tuning, AI Cloud Management, One-Click Auto-Remediation, Unit Economics, Kubernetes Cost Management, Security & Cost, Pre-Flight Impact Analysis, Workload Reliability.
Beyond the schedule: recommendations, rightsizing, and showback
Scheduling is the fastest lever, but it is one of several. ZopNight ships 490 built-in audit rules across AWS (216), GCP (127), and Azure (147) that flag idle, oversized, and orphaned resources, and each recommendation shows the current monthly cost next to the estimated optimized cost so you act on the largest first. 124 of those recommendations are wired to act end to end, 28 one-click and 96 guided: one-click actions run immediately behind an admin-approval gate, and guided actions add a type-to-confirm review so you check the change before it lands. You mark a recommendation applied once you act, or dismiss the ones that do not fit.
Idle detection reads CPU, network, and connection metrics over a rolling window to separate a genuinely idle resource from one with real but intermittent traffic. Rightsizing is guided and computed from measured utilization over a real window, never a flat 24/7 assumption, so the projected figure matches the bill you actually see. For steady-state fleets, VM autoscaling runs in one of three modes derived from the credential’s permissions: monitor, recommend, or autopilot.
What is left after optimization gets attributed rather than hidden. Showback splits shared cost across owning teams and rolls up by cloud tag, GCP label, or Azure tag, with a Sankey cost-flow view that traces spend across provider, account, type, and team and a savings overlay that points straight at the reclaimable flows. A daily anomaly job writes root-cause markers onto the cost trend, an instance resize, a new resource, a reservation expiry, a failed schedule, so a spike explains itself instead of prompting a manual hunt. And 43 read-only tools expose the same data to an AI assistant over MCP, so you can ask an assistant in Claude, Cursor, or Codex for the same numbers.
Best practices that keep the savings
A few habits separate teams that hold onto the savings from teams that watch them drift back:
- Start with non-production and prove it there. Development, staging, QA, and demo environments carry almost no risk and the largest idle share, so they are the right place to build confidence before anyone considers production.
- Schedule by group, not by hand. Bundling an environment into a group like “staging” means one cadence covers every resource in it, and resources you add later inherit the schedule instead of being quietly forgotten.
- Use overrides instead of disabling schedules. When a late deploy needs a box overnight, a time-boxed override with a written reason keeps the schedule intact and expires on its own, so a one-off exception never becomes a permanent leak.
- Watch the audit trail and notifications. Every start, stop, and failure is logged and can post to Slack, Teams, or Google Chat, so a failed action is visible the moment it happens rather than discovered on the next invoice.
- Treat it as an operating rhythm, not a cleanup. The teams that keep the bill down review recommendations on a cadence and let the automation run continuously, instead of a one-off spring-clean that snaps back the moment attention moves on.
Getting started
Getting started is intentionally low-stakes:
- Connect AWS, GCP, and Azure with a read-only role. Nothing is scheduled or changed at this stage.
- Let ZopNight discover your ML Compute and review exactly what it found, filtered by account, region, and status.
- Create a schedule in your timezone and attach the non-production resources or groups you want it to cover.
- Watch the first cycle run, with Slack, Teams, or Google Chat notifications on every start, stop, and failure, then layer in idle cleanup and guided rightsizing.
Production stays excluded by default throughout, and because discovery and recommendations are read-only, you can prove the value before you enable a single action.
Questions we get a lot.
If yours isn't here, email us and we'll answer directly.
How does ZopNight decide what to act on for Azure ML Compute?
It runs built-in audit rules that flag idle, oversized, and orphaned resources, then attaches each finding to your measured usage so you act on the largest first. Scheduling is one-click; rightsizing is guided so you review before anything changes.
Is production at risk?
No. Production is excluded by default; scheduling and rightsizing apply only to the non-production resources you choose.
Does this require changing my infrastructure?
No. Connect a read-only role and enable a schedule. There is no code change, no migration, and no agent to install.
Does stopping a resource delete my data?
No. A scheduled stop preserves attached storage exactly as a normal power-off would; ZopNight stops compute, it never terminates or deletes resources. Your data is intact when the resource starts again.
What access does ZopNight need to begin?
A read-only role. Discovery, cost reporting, and recommendations all run read-only, and ZopNight records a per-action permission verdict so you can see exactly what a credential can and cannot do before you grant anything more.
What happens if a start or stop action fails?
ZopNight retries automatically up to three times, then falls back to a dead-letter queue rather than dropping the action silently. The failure surfaces in the action status and can notify your Slack, Teams, or Google Chat channel.
Which clouds are supported?
AWS, GCP, and Azure from one platform, including Databricks across all three. Schedules, groups, overrides, and recommendations work the same way regardless of provider.