EMR clusters sitting in the WAITING state with no YARN applications for part of the week
What does ZopNight detect here?
ZopNight flags EMR clusters in the `WAITING` state whose measured idle share, from a usage heatmap or the `IsIdle` metric with zero `AppsRunning`, is above 0 and below 95%. An idle cluster still bills per second, so the saving is the cluster cost times that idle share, recovered with an auto-termination policy.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-051 |
| Category | schedule |
| Severity | medium |
| Metric | IsIdle, AppsRunning (AWS/ElasticMapReduce) |
| Threshold | WAITING, idle share between 0% and 95% |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | elasticmapreduce:ListClusters · elasticmapreduce:GetAutoTerminationPolicy · cloudwatch:GetMetricStatistics |
Where it applies
An idle EMR cluster is still a running fleet
EMR on EC2 bills a per-second rate with a one-minute
minimum, and that EMR price is added to the EC2 and EBS price of the servers underneath. When a cluster finishes
its steps and is not set to terminate, it enters the WAITING state, and the
EMR overview says you must
then shut it down yourself. AWS describes the IsIdle metric as indicating a cluster that is no longer
performing work but is still alive and accruing charges.
Finding waiting clusters and their idle time
aws emr list-clusters --cluster-states WAITING \ --query 'Clusters[].[Id,Name,NormalizedInstanceHours]' --output table
aws emr get-auto-termination-policy --cluster-id j-2AXXXXXXGAPLF
aws cloudwatch get-metric-statistics --namespace AWS/ElasticMapReduce --metric-name IsIdle \ --dimensions Name=JobFlowId,Value=j-2AXXXXXXGAPLF \ --start-time 2026-08-26T00:00:00Z --end-time 2026-09-25T00:00:00Z \ --period 3600 --statistics AverageThe average of IsIdle over a period is the share of that period the cluster was idle.
Evidence ZopNight uses
- The cluster’s state is
WAITING, meaning it has no active steps. - ZopNight has a monthly cost for it.
- It has an idle share, preferring a weekly usage heatmap for the cluster and otherwise using the
30-day mean of
IsIdle, providedAppsRunningstayed at zero YARN applications. - The idle share is above 0 and below 95%.
Clusters that fall outside the rule
A cluster with no heatmap and no usable IsIdle series gets no finding. One idle 95% or more of the
time is not a scheduling problem but a likely abandoned cluster, and the rule leaves that call to a
person rather than suggest a recurring timeout. The rule never recommends terminating a cluster on its
own and never shows a fixed monthly figure.
Idle share of the cluster bill
saving = monthly cluster cost x measured idle shareStopping idle time from accruing
- Attach an auto-termination policy so EMR shuts the cluster down after a period of idleness. The
auto-termination docs
allow a timeout from one minute to 7 days, defaulting to one hour:
aws emr put-auto-termination-policy --cluster-id j-2AXXXXXXGAPLF --auto-termination-policy IdleTimeout=3600 - For batch pipelines, launch transient clusters that terminate after their last step.
- For shell scripts or non-YARN work on EMR 6.4.0 or later, touch
/emr/metricscollector/isbusyperiodically so the cluster is not judged idle while working.