Skip to main content
schedule · aws

SageMaker HyperPod clusters under 5% CPU and 10% memory for 7 days that could scale to zero off-hours

resource types
1
rule IDs covered
1
severity
high

What does ZopNight detect here?

ZopNight flags in-service SageMaker HyperPod clusters whose nodes averaged under 5% CPU and 10% memory over 7 days, with CPU never above 20%. The suggested fix is a recurring schedule that scales every instance group to 0 nodes off-hours, and the saving is the cluster cost times its measured idle share.

Signal and threshold

How ZopNight evaluates SageMaker HyperPod clusters under 5% CPU and 10% memory for 7 days that could scale to zero off-hours.
Field Value
Rule IDsRC-1625
Categoryschedule
Severityhigh
Metricnode CPU and memory utilization
ThresholdCPU avg < 5%, memory avg < 10%, CPU peak < 20%
Evaluation window7d
SourceZopNight
Permissions usedsagemaker:ListClusters · sagemaker:DescribeCluster · sagemaker:ListClusterNodes · cloudwatch:GetMetricData

An accelerator cluster that is switched on but doing nothing

HyperPod clusters are built for large training jobs, usually on GPU or Trainium instances, and SageMaker AI pricing lists them by instance hour. Between jobs, a cluster can sit for days with its nodes almost untouched. Because each instance group can be set to zero nodes, per the instance group specification, the cluster definition can stay in place while the expensive hardware is released.

Checking what the cluster is running

Terminal window
aws sagemaker list-clusters --query 'ClusterSummaries[].[ClusterName,ClusterStatus]'
aws sagemaker list-cluster-nodes --cluster-name my-cluster \
--query 'ClusterNodeSummaries[].[InstanceGroupName,InstanceId,InstanceType,LaunchTime]' \
--output table

Compare the node list with your job scheduler: if Slurm has had nothing queued for a week, the nodes are idle.

The near-zero test

  1. The cluster is InService. A cluster already at 0 nodes costs nothing and is not flagged.
  2. Node CPU and memory series both exist for the cluster, each with at least 7 days of data.
  3. Over the 7-day window, average CPU is below 5%, average memory below 10%, and CPU never peaks above 20%, so there are no hidden bursts.
  4. The cluster has a positive monthly cost and is not covered by a Savings Plan or reservation.
  5. ZopNight has a measured weekly usage pattern for the cluster with an idle share above zero.

Clusters left alone

A cluster covered by a Savings Plan or reservation produces no finding, because scaling it down would not stop the committed charge. Clusters orchestrated by Amazon EKS report node metrics to a different place and are not evaluated. With missing metrics, a short data window or no measured idle share, the rule stays silent. Clusters between 5% and 10% CPU are sized, not idle, and are covered by Over-Provisioned SageMaker HyperPod Cluster.

Measured idle share of the cluster bill

Terminal window
saving = monthly cluster cost x measured off-hours idle share

There is no flat percentage behind this number. If ZopNight has no idle measurement, there is no figure and no finding.

Scaling the cluster down when nobody is training

  1. Confirm with the ML team that no training or inference work is scheduled on the cluster.
  2. Review the suggested off-hours window and time zone, then apply it so every instance group scales to 0 nodes off-hours and back before working hours.
  3. To do it by hand, set each group’s InstanceCount to 0 with aws sagemaker update-cluster, keeping a note of the original counts.
  4. If the cluster is not needed at all, delete it instead to release the capacity and quota.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·