Skip to main content
rightsizing · aws

SageMaker batch transform jobs that ran under 10% per-vCPU CPU and 50% memory on their instance type

rule IDs covered
1
severity
low

What does ZopNight detect here?

ZopNight reviews finished SageMaker batch transform jobs whose `CPUUtilization`, divided by the instance vCPU count, averaged under 10% with peaks under 80%, and whose memory averaged under 50%. The finding suggests the next smaller `ml.*` size for future runs and prices the per-run saving from the two instance rates.

Signal and threshold

How ZopNight evaluates SageMaker batch transform jobs that ran under 10% per-vCPU CPU and 50% memory on their instance type.
Field Value
Rule IDsRC-1614
Categoryrightsizing
Severitylow
MetricCPUUtilization, MemoryUtilization (/aws/sagemaker/TransformJobs)
Thresholdper-vCPU CPU avg < 10%, peak < 80%, memory avg < 50%
Evaluation window30d
SourceZopNight
Permissions usedsagemaker:ListTransformJobs · sagemaker:DescribeTransformJob · cloudwatch:GetMetricStatistics

A transform job pays for its instance for the whole run

Batch transform is charged for the instance type you choose, based on the duration of use, according to SageMaker AI pricing. A job that runs on an instance twice the size it needs pays roughly twice as much for the same predictions. Because transform jobs usually run on a schedule, an oversized instance type in the job definition repeats the waste every run.

How to read transform job CPU correctly

The /aws/sagemaker/TransformJobs namespace reports CPUUtilization as the sum across cores: the SageMaker metrics reference gives 0% to 400% as the range on a four-CPU instance. MemoryUtilization is already 0 to 100%. Divide CPU by the vCPU count before comparing it with any threshold.

Terminal window
aws sagemaker list-transform-jobs --status-equals Completed \
--query 'TransformJobSummaries[].[TransformJobName,CreationTime]' --output table
aws sagemaker describe-transform-job --transform-job-name my-transform-job \
--query '[TransformResources.InstanceType,TransformResources.InstanceCount,TransformStartTime,TransformEndTime]'

The job metrics use a Host dimension of the form transform-job-name/instance-id.

What has to hold for a finding

  1. ZopNight resolves the instance’s vCPU count from its ml.* type and divides CPU by it. If the count is unknown, it does not guess.
  2. Normalised average CPU is below 10% and memory averages below 50%, within the 30-day window.
  3. The normalised CPU peak is below 80%, so a job with a heavy phase is not shrunk.
  4. The job has a positive run cost.
  5. A concrete one-size-smaller ml.* type exists and ZopNight has rates for both types.

Jobs that do not get a recommendation

GPU instance types are never downsized by this rule, and nor are the smallest sizes of a family or types outside the ml.* naming. When either rate is missing, or the smaller type is not cheaper, there is no finding. The job has already finished, so the advice applies to its next run; ZopNight does not change the job definition.

Per-run saving from the rate ratio

Terminal window
saving = run cost x (1 - smaller type hourly rate / current type hourly rate)

This assumes the job takes about as long on the smaller instance. A CPU-light job usually does.

Right-sizing the next run

  1. Set a smaller InstanceType in the TransformResources of the next CreateTransformJob call, or in your SDK Transformer configuration.
  2. If the job used several instances, try lowering InstanceCount as well.
  3. Tune MaxConcurrentTransforms, BatchStrategy and MaxPayloadInMB so the smaller instance stays busy. Batch transform caps MaxConcurrentTransforms x MaxPayloadInMB at 100 MB.
  4. Compare the run time and cost of the next run with this one.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·