Skip to main content
schedule · aws

Dev and test SageMaker HyperPod clusters below 50% node CPU with idle hours each week

resource types
1
rule IDs covered
1
severity
medium

What does ZopNight detect here?

ZopNight flags in-service SageMaker HyperPod clusters tagged `env`, `environment` or `stage` as dev, test or staging, or named that way, when node CPU averages under 50% over 30 days and the weekly pattern shows idle hours. A ZopNight schedule scales instance groups to zero off-hours and restores their counts, saving cost times the idle share.

Signal and threshold

How ZopNight evaluates Dev and test SageMaker HyperPod clusters below 50% node CPU with idle hours each week.
Field Value
Rule IDsRC-1612
Categoryschedule
Severitymedium
Metricnode CPU utilization
Thresholddev/test cluster, node CPU avg < 50%, idle share > 0
Evaluation window30d
SourceZopNight
Permissions usedsagemaker:ListClusters · sagemaker:DescribeCluster · sagemaker:ListTags

Development clusters run through nights and weekends by default

A HyperPod cluster bills for every instance in every instance group for as long as those instances exist. The HyperPod overview describes clusters built on thousands of accelerators such as AWS Trainium and NVIDIA GPUs. A cluster used by researchers during the day keeps billing when they go home, unless someone reduces the node counts.

Reviewing environment tags and node counts

Terminal window
aws sagemaker list-clusters --query 'ClusterSummaries[].[ClusterName,ClusterArn]' --output table
aws sagemaker list-tags \
--resource-arn arn:aws:sagemaker:us-east-1:123456789012:cluster/abcd1234efgh
aws sagemaker describe-cluster --cluster-name dev-hyperpod \
--query '[ClusterStatus,InstanceGroups[].[InstanceGroupName,CurrentCount]]'

Signals that must agree

  1. The cluster is InService.
  2. An environment tag (env, environment or stage) says dev, test or staging. Without the tag, ZopNight falls back to the cluster name. A production tag or name blocks the finding, and a cluster with no dev or test signal at all is skipped.
  3. No schedule tag from another tool is present.
  4. Node CPU data covers at least 7 days, and the 30-day average is below 50%. A cluster busy around the clock, or one with no CPU data, is not scheduled on its name alone.
  5. ZopNight has a measured weekly pattern with an idle share above zero, and the cluster is not covered by a Savings Plan or reservation.

Clusters without a schedule suggestion

Clusters orchestrated by Amazon EKS do not supply the node metrics ZopNight reads, so they are skipped. The rule has dropped its old assumption of a fixed 16-hour day plus weekends; without a measured idle share there is no finding. Clusters that are nearly idle all week are handled by Idle SageMaker HyperPod Cluster, Near-Zero Utilization.

Off-hours share of the cluster cost

Terminal window
saving = monthly cluster cost x measured off-hours idle share

Putting the cluster on a schedule

  1. Take the suggested off-hours window and time zone from the recommendation.
  2. Create the schedule in ZopNight. At the stop time it scales every instance group to zero and records the counts; at the start time it restores them.
  3. Tell the team the working window, and ask them to write checkpoints to FSx or S3 so a stop never costs work.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·