Dev and test SageMaker HyperPod clusters below 50% node CPU with idle hours each week
What does ZopNight detect here?
ZopNight flags in-service SageMaker HyperPod clusters tagged `env`, `environment` or `stage` as dev, test or staging, or named that way, when node CPU averages under 50% over 30 days and the weekly pattern shows idle hours. A ZopNight schedule scales instance groups to zero off-hours and restores their counts, saving cost times the idle share.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1612 |
| Category | schedule |
| Severity | medium |
| Metric | node CPU utilization |
| Threshold | dev/test cluster, node CPU avg < 50%, idle share > 0 |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | sagemaker:ListClusters · sagemaker:DescribeCluster · sagemaker:ListTags |
Where it applies
Development clusters run through nights and weekends by default
A HyperPod cluster bills for every instance in every instance group for as long as those instances exist. The HyperPod overview describes clusters built on thousands of accelerators such as AWS Trainium and NVIDIA GPUs. A cluster used by researchers during the day keeps billing when they go home, unless someone reduces the node counts.
Reviewing environment tags and node counts
aws sagemaker list-clusters --query 'ClusterSummaries[].[ClusterName,ClusterArn]' --output table
aws sagemaker list-tags \ --resource-arn arn:aws:sagemaker:us-east-1:123456789012:cluster/abcd1234efgh
aws sagemaker describe-cluster --cluster-name dev-hyperpod \ --query '[ClusterStatus,InstanceGroups[].[InstanceGroupName,CurrentCount]]'Signals that must agree
- The cluster is
InService. - An environment tag (
env,environmentorstage) says dev, test or staging. Without the tag, ZopNight falls back to the cluster name. A production tag or name blocks the finding, and a cluster with no dev or test signal at all is skipped. - No schedule tag from another tool is present.
- Node CPU data covers at least 7 days, and the 30-day average is below 50%. A cluster busy around the clock, or one with no CPU data, is not scheduled on its name alone.
- ZopNight has a measured weekly pattern with an idle share above zero, and the cluster is not covered by a Savings Plan or reservation.
Clusters without a schedule suggestion
Clusters orchestrated by Amazon EKS do not supply the node metrics ZopNight reads, so they are skipped. The rule has dropped its old assumption of a fixed 16-hour day plus weekends; without a measured idle share there is no finding. Clusters that are nearly idle all week are handled by Idle SageMaker HyperPod Cluster, Near-Zero Utilization.
Off-hours share of the cluster cost
saving = monthly cluster cost x measured off-hours idle sharePutting the cluster on a schedule
- Take the suggested off-hours window and time zone from the recommendation.
- Create the schedule in ZopNight. At the stop time it scales every instance group to zero and records the counts; at the start time it restores them.
- Tell the team the working window, and ask them to write checkpoints to FSx or S3 so a stop never costs work.