Completed SageMaker training jobs that ran on-demand instead of Managed Spot Training
What does ZopNight detect here?
ZopNight flags SageMaker training jobs in `Completed` status that ran without `EnableManagedSpotTraining`. The saving applies a flat 70% to the actual billed cost of that run, as an estimate for future runs of the same job; AWS says managed spot training can cut training cost by up to 90% versus on-demand.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1621 |
| Category | discount |
| Severity | low |
| Metric | none — pure configuration read |
| Threshold | Completed job, spot not used |
| Source | ZopNight |
| Permissions used | sagemaker:ListTrainingJobs · sagemaker:DescribeTrainingJob |
Where it applies
Training jobs pay for billable seconds, and Spot shrinks them
Managed Spot Training
runs a training job on EC2 Spot capacity and lets SageMaker handle the interruptions. AWS says it
can optimize training cost by up to 90% compared with on-demand instances. The saving shows up
directly on the job: SageMaker reports both TrainingTimeInSeconds and BillableTimeInSeconds,
and the discount is (1 - BillableTimeInSeconds / TrainingTimeInSeconds) x 100. AWS’s example is a
job that ran 500 seconds and was billed for 100, an 80% saving.
The cost is time. Spot capacity can be interrupted, so jobs can take longer to start or finish, and you set how long SageMaker may wait for capacity.
Finding recent runs that paid full price
aws sagemaker list-training-jobs --status-equals Completed \ --creation-time-after 2026-08-01 --query 'TrainingJobSummaries[].TrainingJobName'
aws sagemaker describe-training-job --training-job-name my-job \ --query '[EnableManagedSpotTraining,TrainingTimeInSeconds,BillableTimeInSeconds,ResourceConfig.InstanceType]'EnableManagedSpotTraining of false, with billable time equal to training time, means the run
paid the on-demand rate for every second.
Which runs count as evidence
Only training jobs are considered.
The job must have finished in Completed status, because a failed, stopped or still-running job
reflects a partial run and would misstate the cost. ZopNight must also know for certain that the
run did not use spot: if that flag is missing or unreadable, nothing fires.
Runs that produce nothing
Jobs that already used managed spot are skipped. So are jobs whose run cost ZopNight cannot find, or that cost nothing: a finding is never shown with a $0 saving. Because the recommendation is about future runs of the job, it does not ask you to change or rerun the completed job itself.
How the saving is estimated
The managed spot discount shows up on the job as fewer billable seconds, so ZopNight applies a fixed estimate to the run’s actual cost:
saving = actual cost of the completed run x 0.70Your realized figure will differ. The job summary after the next spot run gives the real number.
Turning on Managed Spot Training for the next run
- Set
EnableManagedSpotTrainingto true in the training job or estimator configuration. - Set
MaxWaitTimeInSecondsin the stopping condition; the API requires it to be equal to or greater thanMaxRuntimeInSeconds. - Add a
CheckpointConfigwith an S3 URI so an interrupted job resumes from its last checkpoint instead of starting over. - Rerun and compare
BillableTimeInSecondswithTrainingTimeInSeconds.