Single-variant SageMaker endpoints on 2 or more instances running under 10% per-vCPU CPU
What does ZopNight detect here?
ZopNight flags in-service SageMaker real-time endpoints with one production variant and at least 2 instances when `CPUUtilization`, divided by vCPUs, averages under 10% over up to 30 days, with peaks under 80% once 7 days of peaks exist. Removing one of N identical instances saves exactly the endpoint cost divided by N.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1624 |
| Category | rightsizing |
| Severity | medium |
| Metric | CPUUtilization, MemoryUtilization (/aws/sagemaker/Endpoints) |
| Threshold | instances >= 2, per-vCPU CPU avg < 10%, peak < 80% once 7 days of peaks exist |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | sagemaker:ListEndpoints · sagemaker:DescribeEndpoint · cloudwatch:GetMetricStatistics |
Where it applies
Each endpoint instance bills while the endpoint is in service
Real-time inference is charged for the instance type you choose, per SageMaker AI pricing, and every instance behind a production variant counts. An endpoint that was scaled out for a launch and never scaled back pays for the extra instances every hour. If the fleet is identical, each instance is an equal share of the bill.
Checking instance count and load
aws sagemaker list-endpoints --status-equals InService \ --query 'Endpoints[].EndpointName' --output text
aws sagemaker describe-endpoint --endpoint-name my-endpoint \ --query 'ProductionVariants[].[VariantName,CurrentInstanceCount,DesiredInstanceCount]'
aws cloudwatch get-metric-statistics --namespace /aws/sagemaker/Endpoints \ --metric-name CPUUtilization \ --dimensions Name=EndpointName,Value=my-endpoint Name=VariantName,Value=AllTraffic \ --start-time 2026-08-26T00:00:00Z --end-time 2026-09-25T00:00:00Z \ --period 3600 --statistics Average MaximumEndpoint CPUUtilization is the sum across cores, so divide by the instance’s vCPU count.
Readings that trigger the finding
- The endpoint is
InServiceand has 2 or more instances. - It has a single production variant.
- Per-vCPU average CPU over the available window, up to 30 days, is below 10%. Once at least 7 days of peak data exist, the per-vCPU peak must also be below 80%; a younger endpoint is judged on its average alone.
- When memory data is present, average
MemoryUtilizationis below 50%. Missing memory data does not block the finding, but high memory does.
Endpoints left alone
A single-instance endpoint cannot shed an instance; if it serves no traffic at all, see Idle SageMaker Endpoint, No Invocations. Endpoints with more than one production variant are skipped, because one instance type and one utilization series cannot describe several independent variant fleets. If the vCPU count of the instance type is unknown, the CPU cannot be normalised and there is no finding. ZopNight does not change the endpoint itself.
One instance out of N
saving = monthly endpoint cost / instance countThis is exact for a single-variant endpoint, because every instance runs the same type for the same hours. It is also conservative: only one instance is removed per recommendation.
Lowering the instance count
- Lower the desired count in place, which
UpdateEndpointWeightsAndCapacities
supports without a new endpoint configuration:
aws sagemaker update-endpoint-weights-and-capacities --endpoint-name my-endpoint --desired-weights-and-capacities VariantName=AllTraffic,DesiredInstanceCount=2 - If auto scaling manages the variant, lower the scalable target’s minimum capacity as well, so the floor matches the new count.
- Watch
ModelLatencyand error counts for a few days on the smaller fleet.