SageMaker notebooks and endpoints left stopped, failed or out of service for 30 days
What does ZopNight detect here?
ZopNight flags SageMaker notebook instances that have been `Stopped` or `Failed` for at least 30 days, pricing only the ML storage volume that keeps billing, and endpoints stuck `Failed` or `OutOfService` for 30 days that still accrued instance charges. Either way the lever is deletion once nobody needs the resource.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-045 |
| Category | idle |
| Severity | medium |
| Metric | none — pure configuration read |
| Threshold | 30 days in Stopped, Failed or OutOfService |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | sagemaker:ListNotebookInstances · sagemaker:DescribeNotebookInstance · sagemaker:ListEndpoints · sagemaker:DescribeEndpoint |
Stopping a notebook ends compute billing but keeps the disk
The StopNotebookInstance API reference says SageMaker disconnects and preserves the ML storage volume when a notebook instance is stopped, and stops charging for the ML compute instance. The volume itself is not released, so a notebook stopped months ago is still a small storage line on the bill.
Endpoints fail differently. The
DescribeEndpoint status list
says an OutOfService endpoint is not available to take incoming requests, and a Failed one
could not be created, updated or re-scaled, with delete as the only operation left. Yet the object and anything still provisioned behind it stay in the account until someone deletes
it.
Listing stale notebooks and broken endpoints
aws sagemaker list-notebook-instances --status-equals Stopped \ --query 'NotebookInstances[].[NotebookInstanceName,LastModifiedTime]' --output table
aws sagemaker list-notebook-instances --status-equals Failed \ --query 'NotebookInstances[].NotebookInstanceName'
aws sagemaker list-endpoints --status-equals OutOfService \ --query 'Endpoints[].[EndpointName,LastModifiedTime]'LastModifiedTime is only a rough guide to how long a notebook has been stopped, because any
change to the instance resets it.
Two branches, one 30-day rule
Notebooks. The instance must be Stopped or Failed, and it must have been in that status for
at least 30 days. ZopNight measures that from its own record of when the status changed. For a
notebook stopped directly in the console, which ZopNight never saw happen, it falls back to the
LastModifiedTime AWS reports. That timestamp can only understate how long the notebook has been
stopped, so the fallback can delay a finding but never rush one. The volume size must also be
known.
Endpoints. The endpoint must be Failed or OutOfService and have stayed that way for at least
30 days in ZopNight’s record, and it must carry a positive cost for the period.
What keeps a resource off the list
A notebook or endpoint in any other status is ignored. With no record of when the status began and no AWS timestamp to fall back on, the 30 days are unproven and there is no finding. A notebook with no discovered volume size is skipped instead of being priced at a guessed default, and an endpoint that accrued no instance hours in the period has nothing to recover, so it is skipped too.
Two different savings figures
notebook saving = volume size in GB x SageMaker ML storage rate per GB-monthendpoint saving = endpoint monthly costcost after fix = 0The notebook figure is deliberately small: it is the storage that leaked, not the compute price of a machine that is no longer running.
Cleaning up
- For a notebook, ask its owner whether anything on the volume is still needed and copy it to S3 if so.
- For an endpoint, check
InvocationsPerInstancehistory to confirm nothing depended on it. - Delete with
aws sagemaker delete-notebook-instance --notebook-instance-name my-notebook(a notebook must be stopped first) oraws sagemaker delete-endpoint --endpoint-name my-endpoint. - For ongoing notebook work, consider SageMaker AI Studio instead of standalone notebook instances.