# Idle SageMaker Resource

> Flags SageMaker notebooks stopped or failed for 30+ days (storage residual) and endpoints failed or out of service for 30+ days.

Source: https://zop.dev/integrations/aws/recommendations/idle-sagemaker-resource

---

## Stopping a notebook ends compute billing but keeps the disk

The [StopNotebookInstance API reference](https://docs.aws.amazon.com/sagemaker/latest/APIReference/API_StopNotebookInstance.html)
says SageMaker disconnects and preserves the ML storage volume when a notebook instance is
stopped, and stops charging for the ML compute instance. The volume itself is not released, so a
notebook stopped months ago is still a small storage line on the bill.

Endpoints fail differently. The
[DescribeEndpoint status list](https://docs.aws.amazon.com/sagemaker/latest/APIReference/API_DescribeEndpoint.html)
says an `OutOfService` endpoint is not available to take incoming requests, and a `Failed` one
could not be created, updated or re-scaled, with delete as the only operation left. Yet the object and anything still provisioned behind it stay in the account until someone deletes
it.

## Listing stale notebooks and broken endpoints

```bash
aws sagemaker list-notebook-instances --status-equals Stopped \
  --query 'NotebookInstances[].[NotebookInstanceName,LastModifiedTime]' --output table

aws sagemaker list-notebook-instances --status-equals Failed \
  --query 'NotebookInstances[].NotebookInstanceName'

aws sagemaker list-endpoints --status-equals OutOfService \
  --query 'Endpoints[].[EndpointName,LastModifiedTime]'
```

`LastModifiedTime` is only a rough guide to how long a notebook has been stopped, because any
change to the instance resets it.

## Two branches, one 30-day rule

**Notebooks.** The instance must be `Stopped` or `Failed`, and it must have been in that status for
at least 30 days. ZopNight measures that from its own record of when the status changed. For a
notebook stopped directly in the console, which ZopNight never saw happen, it falls back to the
`LastModifiedTime` AWS reports. That timestamp can only understate how long the notebook has been
stopped, so the fallback can delay a finding but never rush one. The volume size must also be
known.

**Endpoints.** The endpoint must be `Failed` or `OutOfService` and have stayed that way for at least
30 days in ZopNight's record, and it must carry a positive cost for the period.

## What keeps a resource off the list

A notebook or endpoint in any other status is ignored. With no record of when the status began and
no AWS timestamp to fall back on, the 30 days are unproven and there is no finding. A notebook with
no discovered volume size is skipped instead of being priced at a guessed default, and an endpoint
that accrued no instance hours in the period has nothing to recover, so it is skipped too.

## Two different savings figures

```text
notebook saving  = volume size in GB x SageMaker ML storage rate per GB-month
endpoint saving  = endpoint monthly cost
cost after fix   = 0
```

The notebook figure is deliberately small: it is the storage that leaked, not the compute price of
a machine that is no longer running.

## Cleaning up

1. For a notebook, ask its owner whether anything on the volume is still needed and copy it to S3
   if so.
2. For an endpoint, check `InvocationsPerInstance` history to confirm nothing depended on it.
3. Delete with `aws sagemaker delete-notebook-instance --notebook-instance-name my-notebook` (a
   notebook must be stopped first) or
   `aws sagemaker delete-endpoint --endpoint-name my-endpoint`.
4. For ongoing notebook work, consider SageMaker AI Studio instead of standalone notebook
   instances.

**Warning**
Deleting a notebook instance deletes its ML storage volume and everything on it. Copy the
files out first.
