Idle SageMaker HyperPod Cluster, Near-Zero Utilisation
A HyperPod cluster holds its accelerator capacity reserved for as long as it exists, and reserved accelerators are the most expensive idle resource in the account, nothing else wastes at this rate.
Free to start. No card. The playground just needs your work email.
AWS
Found, explained, handed over.
Findings with a dollar figure attached, idle, oversized, orphaned, unscheduled and undiscounted spend.
Detect
ZopNight checks this automatically across AWS, with read-only access to the account.
Explain
Every finding says exactly what to change: Scale down the idle HyperPod cluster. Where the saving can be proven it is priced; where it cannot, the finding says so.
Fix
The finding opens with the fix already written out, step by step, so it is one ticket, not an investigation.
- Applies to
- SageMaker Clusters on AWS
- The fix, by hand
- Confirm no training/inference workload is scheduled on this cluster.
- Review the recommended off-hours schedule (start/stop cron, timezone).
- Apply it to scale every instance group to 0 nodes during off-hours (the HyperPod off state) and back up before working hours.
- If the cluster is no longer needed at all, delete it instead to free the reservation/quota.
These are the steps the finding carries in the product.
- Category
- Scheduling. Resources running 24x7 with a clear off-hours usage pattern, ready for start and stop scheduling.
- Where it appears
- The Savings tab of Recommendations, with every affected resource listed.
- Rule ID
RC-1625- Full reference
Related checks
See the always-on resources in your account.
Connect a read-only role and the first pass runs on your own estate. This check, and the rest of the catalogue, with it.
Prefer to talk it through first? Book 20 minutes with the team.
- $30M+annualised cloud spend under management
- 550K+resources tracked since launch
- 20-60%off the bill in the first month
- SOC 2Type II report, plus ISO 27001
Figures published on zop.dev.