SageMaker HyperPod clusters with two or more nodes running at 5 to 10% CPU and under 50% memory
What does ZopNight detect here?
ZopNight flags an in-service SageMaker HyperPod cluster of at least 2 nodes when 30 days of node metrics show CPU averaging between 5% and 10%, memory under 50% and CPU peaks under 80%. The suggested change is removing one node, and the saving is roughly the cluster cost divided by its node count.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1626 |
| Category | rightsizing |
| Severity | medium |
| Metric | node CPU and memory utilization |
| Threshold | CPU avg 5-10%, memory avg < 50%, CPU peak < 80% |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | sagemaker:ListClusters · sagemaker:DescribeCluster · cloudwatch:GetMetricData |
Where it applies
Every HyperPod node bills for the hour, busy or not
SageMaker HyperPod clusters are priced per instance hour, per the instance tables on SageMaker AI pricing, and that price does not cover connected services such as Amazon EKS, FSx for Lustre or S3. A node that runs at a few percent CPU all month costs the same as one saturated by a training job. HyperPod lets you set the node count of each instance group separately, so an oversized group can shrink without rebuilding the cluster.
Looking at node counts and utilization
aws sagemaker list-clusters \ --query 'ClusterSummaries[].[ClusterName,ClusterStatus]' --output table
aws sagemaker describe-cluster --cluster-name my-cluster \ --query 'InstanceGroups[].[InstanceGroupName,InstanceType,CurrentCount,TargetCount]' \ --output tableThen open the cluster’s node CPU and memory graphs in CloudWatch for the past 30 days and compare each instance group. The group whose nodes sit lowest is the one to trim.
Four readings that must all line up
- The cluster is
InServiceand has 2 or more nodes. A single-node cluster cannot lose a node; if it is idle, the case belongs to Idle SageMaker HyperPod Cluster, Near-Zero Utilization. - Average node CPU over 30 days is at least 5% and below 10%. The 5% floor keeps this rule and the idle rule from firing on the same cluster.
- Average node memory is below 50%.
- The trusted CPU peak is below 80%, so a cluster that bursts hard during jobs is left alone.
Clusters this rule leaves alone
Clusters with fewer than 2 nodes, or whose node count is unknown, are skipped. Missing CPU or memory series mean no finding, and so does a cluster with no monthly cost. ZopNight receives node metrics for Slurm-orchestrated clusters; EKS-orchestrated clusters report node metrics elsewhere, so they are not evaluated by this rule. The rule is advisory: ZopNight does not resize HyperPod clusters itself.
Estimating the value of one fewer node
saving = monthly cluster cost / node countThis assumes all nodes cost the same. A cluster that mixes instance groups of different types will save more or less depending on which group you shrink, and the recommendation text says so.
Trimming the right instance group
- Pick the instance group whose nodes are least used, not simply the most expensive one.
- Lower that group’s
InstanceCountwithaws sagemaker update-cluster. The instance group specification in the request must still carry the group’s required settings, such as its execution role. - Watch job queue wait times and training throughput for a week on the smaller fleet.
- Repeat one node at a time if utilization stays low.