Skip to main content
discount · aws

EMR clusters whose task nodes run On-Demand instead of Spot

resource types
1
rule IDs covered
1
severity
low

What does ZopNight detect here?

ZopNight flags Amazon EMR clusters whose task instance group or fleet uses the `ON_DEMAND` market, when the task instance type and node count are known. The saving is the On-Demand minus Spot rate for that type, times the task node count and the hours the cluster runs; primary and core nodes are never suggested for Spot.

Signal and threshold

How ZopNight evaluates EMR clusters whose task nodes run On-Demand instead of Spot.
Field Value
Rule IDsRC-052
Categorydiscount
Severitylow
Metricnone — pure configuration read
Thresholdtask market ON_DEMAND, 1+ task node
SourceZopNight
Permissions usedelasticmapreduce:ListClusters · elasticmapreduce:DescribeCluster · elasticmapreduce:ListInstances · elasticmapreduce:ListInstanceFleets · elasticmapreduce:ListInstanceGroups

Task nodes are the part of an EMR cluster built for Spot

An EMR cluster has primary, core and task nodes. Core nodes hold HDFS data; task nodes only run work, so losing one costs a retry rather than data. AWS’s instance and Spot guidelines treat task nodes as the usual home for Spot Instances, and EMR can keep application master processes on core nodes so jobs survive a task node being reclaimed.

One caveat from the same guide: from the EMR 6.x release series, the YARN node labels feature is disabled by default, so application masters may run on task nodes unless you turn it on. Check that before moving a long-running job’s task capacity to Spot.

Seeing how the task capacity is bought

Terminal window
aws emr list-clusters --active
aws emr describe-cluster --cluster-id j-EXAMPLE \
--query 'Cluster.[InstanceCollectionType,ReleaseLabel]'
aws emr list-instances --cluster-id j-EXAMPLE --instance-group-types TASK \
--instance-states RUNNING --query 'Instances[].[InstanceType,Market]'

Market is ON_DEMAND or SPOT. For clusters built from instance fleets, use --instance-fleet-type TASK instead, or aws emr list-instance-fleets to see the On-Demand and Spot target capacities.

What ZopNight needs to know about the task nodes

  • The cluster’s task capacity is marked ON_DEMAND.
  • A single task instance type and a task node count of at least one can be attributed. Fleets that mix several instance specs, or weight capacity above one unit per instance, are not attributed and stay unpriced.
  • The cluster’s priced instance type matches that task instance type, so the rates belong to the right SKU.
  • A positive monthly cost and live On-Demand and Spot rates for the task instance type.

Clusters left out

A cluster with only primary and core nodes produces nothing; recommending Spot for HDFS nodes is not something this rule does. Clusters already using Spot for tasks are skipped. Missing or implausible Spot rates end the evaluation without a finding. A cluster that has stopped doing work is a different problem, covered by Idle EMR Cluster.

Pricing only the task nodes

Terminal window
hours = 730 x measured uptime (730 when uptime is unknown)
saving = (On-Demand rate - Spot rate) x task node count x hours
cap = monthly cluster cost x Spot discount fraction

The cap stops the task-node estimate from ever claiming more than the whole cluster could save.

Moving task capacity to Spot

AWS does not let you change a purchasing option while the cluster runs. For task nodes you add new capacity and remove the old:

  1. Add a Spot task group: aws emr add-instance-groups --cluster-id j-EXAMPLE --instance-groups InstanceGroupType=TASK,InstanceType=m5.xlarge,InstanceCount=4,Market=SPOT
  2. For fleets, set TargetSpotCapacity with aws emr modify-instance-fleet --cluster-id j-EXAMPLE --instance-fleet InstanceFleetId=if-EXAMPLE,TargetSpotCapacity=4,TargetOnDemandCapacity=0
  3. List several instance types in a fleet (up to five in the console, up to 30 through the CLI or API with an allocation strategy) to widen Spot capacity.
  4. Resize the old On-Demand task group to zero once jobs run cleanly.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·