Skip to main content
rightsizing · aws

SageMaker HyperPod clusters with two or more nodes running at 5 to 10% CPU and under 50% memory

resource types
1
rule IDs covered
1
severity
medium

What does ZopNight detect here?

ZopNight flags an in-service SageMaker HyperPod cluster of at least 2 nodes when 30 days of node metrics show CPU averaging between 5% and 10%, memory under 50% and CPU peaks under 80%. The suggested change is removing one node, and the saving is roughly the cluster cost divided by its node count.

Signal and threshold

How ZopNight evaluates SageMaker HyperPod clusters with two or more nodes running at 5 to 10% CPU and under 50% memory.
Field Value
Rule IDsRC-1626
Categoryrightsizing
Severitymedium
Metricnode CPU and memory utilization
ThresholdCPU avg 5-10%, memory avg < 50%, CPU peak < 80%
Evaluation window30d
SourceZopNight
Permissions usedsagemaker:ListClusters · sagemaker:DescribeCluster · cloudwatch:GetMetricData

Every HyperPod node bills for the hour, busy or not

SageMaker HyperPod clusters are priced per instance hour, per the instance tables on SageMaker AI pricing, and that price does not cover connected services such as Amazon EKS, FSx for Lustre or S3. A node that runs at a few percent CPU all month costs the same as one saturated by a training job. HyperPod lets you set the node count of each instance group separately, so an oversized group can shrink without rebuilding the cluster.

Looking at node counts and utilization

Terminal window
aws sagemaker list-clusters \
--query 'ClusterSummaries[].[ClusterName,ClusterStatus]' --output table
aws sagemaker describe-cluster --cluster-name my-cluster \
--query 'InstanceGroups[].[InstanceGroupName,InstanceType,CurrentCount,TargetCount]' \
--output table

Then open the cluster’s node CPU and memory graphs in CloudWatch for the past 30 days and compare each instance group. The group whose nodes sit lowest is the one to trim.

Four readings that must all line up

  • The cluster is InService and has 2 or more nodes. A single-node cluster cannot lose a node; if it is idle, the case belongs to Idle SageMaker HyperPod Cluster, Near-Zero Utilization.
  • Average node CPU over 30 days is at least 5% and below 10%. The 5% floor keeps this rule and the idle rule from firing on the same cluster.
  • Average node memory is below 50%.
  • The trusted CPU peak is below 80%, so a cluster that bursts hard during jobs is left alone.

Clusters this rule leaves alone

Clusters with fewer than 2 nodes, or whose node count is unknown, are skipped. Missing CPU or memory series mean no finding, and so does a cluster with no monthly cost. ZopNight receives node metrics for Slurm-orchestrated clusters; EKS-orchestrated clusters report node metrics elsewhere, so they are not evaluated by this rule. The rule is advisory: ZopNight does not resize HyperPod clusters itself.

Estimating the value of one fewer node

Terminal window
saving = monthly cluster cost / node count

This assumes all nodes cost the same. A cluster that mixes instance groups of different types will save more or less depending on which group you shrink, and the recommendation text says so.

Trimming the right instance group

  1. Pick the instance group whose nodes are least used, not simply the most expensive one.
  2. Lower that group’s InstanceCount with aws sagemaker update-cluster. The instance group specification in the request must still carry the group’s required settings, such as its execution role.
  3. Watch job queue wait times and training throughput for a week on the smaller fleet.
  4. Repeat one node at a time if utilization stays low.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·