EMR clusters whose task nodes run On-Demand instead of Spot
What does ZopNight detect here?
ZopNight flags Amazon EMR clusters whose task instance group or fleet uses the `ON_DEMAND` market, when the task instance type and node count are known. The saving is the On-Demand minus Spot rate for that type, times the task node count and the hours the cluster runs; primary and core nodes are never suggested for Spot.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-052 |
| Category | discount |
| Severity | low |
| Metric | none — pure configuration read |
| Threshold | task market ON_DEMAND, 1+ task node |
| Source | ZopNight |
| Permissions used | elasticmapreduce:ListClusters · elasticmapreduce:DescribeCluster · elasticmapreduce:ListInstances · elasticmapreduce:ListInstanceFleets · elasticmapreduce:ListInstanceGroups |
Where it applies
Task nodes are the part of an EMR cluster built for Spot
An EMR cluster has primary, core and task nodes. Core nodes hold HDFS data; task nodes only run work, so losing one costs a retry rather than data. AWS’s instance and Spot guidelines treat task nodes as the usual home for Spot Instances, and EMR can keep application master processes on core nodes so jobs survive a task node being reclaimed.
One caveat from the same guide: from the EMR 6.x release series, the YARN node labels feature is disabled by default, so application masters may run on task nodes unless you turn it on. Check that before moving a long-running job’s task capacity to Spot.
Seeing how the task capacity is bought
aws emr list-clusters --active
aws emr describe-cluster --cluster-id j-EXAMPLE \ --query 'Cluster.[InstanceCollectionType,ReleaseLabel]'
aws emr list-instances --cluster-id j-EXAMPLE --instance-group-types TASK \ --instance-states RUNNING --query 'Instances[].[InstanceType,Market]'Market is ON_DEMAND or SPOT. For clusters built from instance fleets, use
--instance-fleet-type TASK instead, or aws emr list-instance-fleets to see the On-Demand and
Spot target capacities.
What ZopNight needs to know about the task nodes
- The cluster’s task capacity is marked
ON_DEMAND. - A single task instance type and a task node count of at least one can be attributed. Fleets that mix several instance specs, or weight capacity above one unit per instance, are not attributed and stay unpriced.
- The cluster’s priced instance type matches that task instance type, so the rates belong to the right SKU.
- A positive monthly cost and live On-Demand and Spot rates for the task instance type.
Clusters left out
A cluster with only primary and core nodes produces nothing; recommending Spot for HDFS nodes is not something this rule does. Clusters already using Spot for tasks are skipped. Missing or implausible Spot rates end the evaluation without a finding. A cluster that has stopped doing work is a different problem, covered by Idle EMR Cluster.
Pricing only the task nodes
hours = 730 x measured uptime (730 when uptime is unknown)saving = (On-Demand rate - Spot rate) x task node count x hourscap = monthly cluster cost x Spot discount fractionThe cap stops the task-node estimate from ever claiming more than the whole cluster could save.
Moving task capacity to Spot
AWS does not let you change a purchasing option while the cluster runs. For task nodes you add new capacity and remove the old:
- Add a Spot task group:
aws emr add-instance-groups --cluster-id j-EXAMPLE --instance-groups InstanceGroupType=TASK,InstanceType=m5.xlarge,InstanceCount=4,Market=SPOT - For fleets, set
TargetSpotCapacitywithaws emr modify-instance-fleet --cluster-id j-EXAMPLE --instance-fleet InstanceFleetId=if-EXAMPLE,TargetSpotCapacity=4,TargetOnDemandCapacity=0 - List several instance types in a fleet (up to five in the console, up to 30 through the CLI or API with an allocation strategy) to widen Spot capacity.
- Resize the old On-Demand task group to zero once jobs run cleanly.