SageMaker HyperPod clusters under 5% CPU and 10% memory for 7 days that could scale to zero off-hours
What does ZopNight detect here?
ZopNight flags in-service SageMaker HyperPod clusters whose nodes averaged under 5% CPU and 10% memory over 7 days, with CPU never above 20%. The suggested fix is a recurring schedule that scales every instance group to 0 nodes off-hours, and the saving is the cluster cost times its measured idle share.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1625 |
| Category | schedule |
| Severity | high |
| Metric | node CPU and memory utilization |
| Threshold | CPU avg < 5%, memory avg < 10%, CPU peak < 20% |
| Evaluation window | 7d |
| Source | ZopNight |
| Permissions used | sagemaker:ListClusters · sagemaker:DescribeCluster · sagemaker:ListClusterNodes · cloudwatch:GetMetricData |
Where it applies
An accelerator cluster that is switched on but doing nothing
HyperPod clusters are built for large training jobs, usually on GPU or Trainium instances, and SageMaker AI pricing lists them by instance hour. Between jobs, a cluster can sit for days with its nodes almost untouched. Because each instance group can be set to zero nodes, per the instance group specification, the cluster definition can stay in place while the expensive hardware is released.
Checking what the cluster is running
aws sagemaker list-clusters --query 'ClusterSummaries[].[ClusterName,ClusterStatus]'
aws sagemaker list-cluster-nodes --cluster-name my-cluster \ --query 'ClusterNodeSummaries[].[InstanceGroupName,InstanceId,InstanceType,LaunchTime]' \ --output tableCompare the node list with your job scheduler: if Slurm has had nothing queued for a week, the nodes are idle.
The near-zero test
- The cluster is
InService. A cluster already at 0 nodes costs nothing and is not flagged. - Node CPU and memory series both exist for the cluster, each with at least 7 days of data.
- Over the 7-day window, average CPU is below 5%, average memory below 10%, and CPU never peaks above 20%, so there are no hidden bursts.
- The cluster has a positive monthly cost and is not covered by a Savings Plan or reservation.
- ZopNight has a measured weekly usage pattern for the cluster with an idle share above zero.
Clusters left alone
A cluster covered by a Savings Plan or reservation produces no finding, because scaling it down would not stop the committed charge. Clusters orchestrated by Amazon EKS report node metrics to a different place and are not evaluated. With missing metrics, a short data window or no measured idle share, the rule stays silent. Clusters between 5% and 10% CPU are sized, not idle, and are covered by Over-Provisioned SageMaker HyperPod Cluster.
Measured idle share of the cluster bill
saving = monthly cluster cost x measured off-hours idle shareThere is no flat percentage behind this number. If ZopNight has no idle measurement, there is no figure and no finding.
Scaling the cluster down when nobody is training
- Confirm with the ML team that no training or inference work is scheduled on the cluster.
- Review the suggested off-hours window and time zone, then apply it so every instance group scales to 0 nodes off-hours and back before working hours.
- To do it by hand, set each group’s
InstanceCountto 0 withaws sagemaker update-cluster, keeping a note of the original counts. - If the cluster is not needed at all, delete it instead to release the capacity and quota.