SageMaker HyperPod Cluster Non-Production Scheduling Opportunity
The cluster holds accelerator capacity reserved around the clock, and a non-production cluster is reserving it for the hours nobody trains.
Free to start. No card. The playground just needs your work email.
AWS
Found, explained, handed over.
Findings with a dollar figure attached, idle, oversized, orphaned, unscheduled and undiscounted spend.
Detect
ZopNight checks this automatically across AWS, with read-only access to the account.
Explain
Every finding says exactly what to change: Schedule the non-production HyperPod cluster off-hours. Where the saving can be proven it is priced; where it cannot, the finding says so.
Fix
The finding opens with the fix already written out, step by step, so it is one ticket, not an investigation.
- Applies to
- SageMaker Clusters on AWS
- The fix, by hand
- Recommended off-hours window (start/stop cron, timezone) from the heatmap.
- Create a schedule in ZopNight to scale all instance groups to zero during off-hours; the executor's aws-sagemaker-cluster provider restores the saved counts on start.
- Automate the scale-down/scale-up on the cron windows above.
These are the steps the finding carries in the product.
- Category
- Scheduling. Resources running 24x7 with a clear off-hours usage pattern, ready for start and stop scheduling.
- Where it appears
- The Savings tab of Recommendations, with every affected resource listed.
- Rule ID
RC-1612- Full reference
SageMaker HyperPod Cluster Non-Production Scheduling Opportunity
Related checks
See the always-on resources in your account.
Connect a read-only role and the first pass runs on your own estate. This check, and the rest of the catalogue, with it.
Prefer to talk it through first? Book 20 minutes with the team.
- $30M+annualised cloud spend under management
- 550K+resources tracked since launch
- 20-60%off the bill in the first month
- SOC 2Type II report, plus ISO 27001
Figures published on zop.dev.