Skip to main content
schedule · aws

EMR clusters sitting in the WAITING state with no YARN applications for part of the week

resource types
1
rule IDs covered
1
severity
medium

What does ZopNight detect here?

ZopNight flags EMR clusters in the `WAITING` state whose measured idle share, from a usage heatmap or the `IsIdle` metric with zero `AppsRunning`, is above 0 and below 95%. An idle cluster still bills per second, so the saving is the cluster cost times that idle share, recovered with an auto-termination policy.

Signal and threshold

How ZopNight evaluates EMR clusters sitting in the WAITING state with no YARN applications for part of the week.
Field Value
Rule IDsRC-051
Categoryschedule
Severitymedium
MetricIsIdle, AppsRunning (AWS/ElasticMapReduce)
ThresholdWAITING, idle share between 0% and 95%
Evaluation window30d
SourceZopNight
Permissions usedelasticmapreduce:ListClusters · elasticmapreduce:GetAutoTerminationPolicy · cloudwatch:GetMetricStatistics

An idle EMR cluster is still a running fleet

EMR on EC2 bills a per-second rate with a one-minute minimum, and that EMR price is added to the EC2 and EBS price of the servers underneath. When a cluster finishes its steps and is not set to terminate, it enters the WAITING state, and the EMR overview says you must then shut it down yourself. AWS describes the IsIdle metric as indicating a cluster that is no longer performing work but is still alive and accruing charges.

Finding waiting clusters and their idle time

Terminal window
aws emr list-clusters --cluster-states WAITING \
--query 'Clusters[].[Id,Name,NormalizedInstanceHours]' --output table
aws emr get-auto-termination-policy --cluster-id j-2AXXXXXXGAPLF
aws cloudwatch get-metric-statistics --namespace AWS/ElasticMapReduce --metric-name IsIdle \
--dimensions Name=JobFlowId,Value=j-2AXXXXXXGAPLF \
--start-time 2026-08-26T00:00:00Z --end-time 2026-09-25T00:00:00Z \
--period 3600 --statistics Average

The average of IsIdle over a period is the share of that period the cluster was idle.

Evidence ZopNight uses

  1. The cluster’s state is WAITING, meaning it has no active steps.
  2. ZopNight has a monthly cost for it.
  3. It has an idle share, preferring a weekly usage heatmap for the cluster and otherwise using the 30-day mean of IsIdle, provided AppsRunning stayed at zero YARN applications.
  4. The idle share is above 0 and below 95%.

Clusters that fall outside the rule

A cluster with no heatmap and no usable IsIdle series gets no finding. One idle 95% or more of the time is not a scheduling problem but a likely abandoned cluster, and the rule leaves that call to a person rather than suggest a recurring timeout. The rule never recommends terminating a cluster on its own and never shows a fixed monthly figure.

Idle share of the cluster bill

Terminal window
saving = monthly cluster cost x measured idle share

Stopping idle time from accruing

  1. Attach an auto-termination policy so EMR shuts the cluster down after a period of idleness. The auto-termination docs allow a timeout from one minute to 7 days, defaulting to one hour: aws emr put-auto-termination-policy --cluster-id j-2AXXXXXXGAPLF --auto-termination-policy IdleTimeout=3600
  2. For batch pipelines, launch transient clusters that terminate after their last step.
  3. For shell scripts or non-YARN work on EMR 6.4.0 or later, touch /emr/metricscollector/isbusy periodically so the cluster is not judged idle while working.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·