SageMaker Endpoint Instance Count Over-Provisioned
An endpoint bills for every instance behind it continuously, so instances beyond what the invocation rate needs are billed idle.
Free to start. No card. The playground just needs your work email.
AWS
Found, explained, handed over.
Findings with a dollar figure attached, idle, oversized, orphaned, unscheduled and undiscounted spend.
Detect
ZopNight checks this automatically across AWS, with read-only access to the account.
Explain
Every finding says exactly what to change: Reduce the SageMaker endpoint's instance count. Where the saving can be proven it is priced; where it cannot, the finding says so.
Fix
The finding opens with the fix already written out, step by step, so it is one ticket, not an investigation.
- Applies to
- SageMaker Endpoints on AWS
- The fix, by hand
- Review the CloudWatch CPUUtilization/MemoryUtilization for this endpoint's production variant.
- Lower DesiredInstanceCount in UpdateEndpointWeightsAndCapacities (or reduce InitialInstanceCount in the endpoint config).
- If autoscaling is enabled, lower the MinCapacity on the scalable target.
- Confirm utilisation stays healthy and latency is within SLA on the smaller fleet.
These are the steps the finding carries in the product.
- Category
- Rightsizing. Resources sized for headroom they never use. A smaller size runs the same workload for less.
- Where it appears
- The Savings tab of Recommendations, with every affected resource listed.
- Rule ID
RC-1624- Full reference
Related checks
See the oversized resources in your account.
Connect a read-only role and the first pass runs on your own estate. This check, and the rest of the catalogue, with it.
Prefer to talk it through first? Book 20 minutes with the team.
- $30M+annualised cloud spend under management
- 550K+resources tracked since launch
- 20-60%off the bill in the first month
- SOC 2Type II report, plus ISO 27001
Figures published on zop.dev.