SageMaker endpoints averaging at most one invocation a day that could run on a smaller instance
What does ZopNight detect here?
ZopNight flags in-service SageMaker real-time endpoints whose `Invocations` metric shows traffic above zero but no more than 1 request a day on average over the last 7 days. The finding proposes the next smaller `ml.*` instance and prices it from the two hourly rates, skipping endpoints covered by a Savings Plan.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1501 |
| Category | rightsizing |
| Severity | medium |
| Metric | Invocations, InvocationsPerInstance |
| Threshold | 0 < invocations per day <= 1 |
| Evaluation window | 7d |
| Source | ZopNight |
| Permissions used | sagemaker:ListEndpoints · sagemaker:DescribeEndpointConfig · cloudwatch:GetMetricStatistics |
Where it applies
A trickle of requests on a full-size instance
A real-time endpoint keeps its instances running whether it receives a thousand requests a minute or
one a day. The Invocations metric in the AWS/SageMaker namespace counts InvokeEndpoint
requests, and the
metrics reference
says to read it with the Sum statistic. An endpoint whose daily sum is in single digits is running
an instance sized for traffic that is not there.
Counting invocations per day
aws cloudwatch get-metric-statistics --namespace AWS/SageMaker --metric-name Invocations \ --dimensions Name=EndpointName,Value=my-endpoint Name=VariantName,Value=AllTraffic \ --start-time 2026-08-26T00:00:00Z --end-time 2026-09-25T00:00:00Z \ --period 86400 --statistics Sum
aws sagemaker describe-endpoint-config --endpoint-config-name my-endpoint-config \ --query 'ProductionVariants[].[VariantName,InstanceType,InitialInstanceCount]'The low-but-not-zero band
- The endpoint is
InServiceand not covered by a Savings Plan or reservation, since on-demand rate differences overstate savings on committed capacity. - At least 7 days of invocation data exist in the metric history ZopNight holds.
- Total invocations over the last 7 days, divided by 7, is above 0 and at most 1.
- A concrete one-size-smaller
ml.*instance exists, and ZopNight has hourly rates for both types. - The endpoint has a positive monthly cost.
If invocation data is missing or covers fewer than 7 days, ZopNight accepts an
invocations_low=true tag on the endpoint in place of steps 2 and 3.
Endpoints outside the band
An endpoint with exactly zero invocations is not a downsize case; it belongs to
Idle SageMaker Endpoint, No Invocations,
and the lower bound keeps the two rules from giving opposite advice on one endpoint. GPU instance
types, the smallest size of a family, and types outside ml.* naming produce nothing. The same
applies when a rate is missing or the smaller type is not cheaper.
Rate-ratio saving
saving = monthly endpoint cost x (1 - smaller type rate / current type rate)There is no fallback percentage: without both rates there is no figure and no finding.
Options for a rarely used endpoint
- Resize the variant to the smaller instance by creating a new endpoint configuration and calling
aws sagemaker update-endpoint. ZopNight can also apply this resize from the recommendation. - For traffic this sparse, consider Serverless Inference, which AWS describes as suited to workloads with idle periods between traffic spurts that can tolerate cold starts.
- Check
ModelLatencyafter the change to confirm the smaller instance still meets its target.