Skip to main content
idle · gcp

Long-lived Dataproc clusters that ran no YARN applications for 30 days

resource types
1
rule IDs covered
1
severity
medium

What does ZopNight detect here?

Dataproc clusters are flagged when active YARN applications, from `dataproc.googleapis.com/cluster/yarn/apps`, stay at or below 0.5 on average and at peak for 30 days. The cluster's Compute Engine VMs and the Dataproc fee of $0.010 per vCPU-hour keep billing regardless, so ZopNight counts the whole cluster cost as the saving.

Signal and threshold

How ZopNight evaluates Long-lived Dataproc clusters that ran no YARN applications for 30 days.
Field Value
Rule IDsRC-1210
Categoryidle
Severitymedium
Metricdataproc.googleapis.com/cluster/yarn/apps
ThresholdYARN active applications at or below 0.5
Evaluation window30d
SourceZopNight
Permissions useddataproc.clusters.list · monitoring.timeSeries.list

What an idle Dataproc cluster keeps paying for

A Dataproc cluster on Compute Engine is billed on uptime, not on work. Google’s Dataproc pricing charges a management fee of $0.010 per vCPU per hour across master, worker and secondary worker nodes, billed by the second, and that fee is in addition to the Compute Engine price of every VM in the cluster. Persistent disks attached to the nodes bill too.

Clusters built for a one-off analysis and never torn down are the common case. They look harmless in the console and cost as much as a busy cluster of the same size.

Checking for YARN activity

Terminal window
gcloud dataproc clusters list --region=REGION

In Metrics Explorer, chart dataproc.googleapis.com/cluster/yarn/apps for the cluster over 30 days. The metric counts active YARN applications, with a status label for running, pending and other states. Zero throughout means no Spark, Hive or MapReduce job ran.

What ZopNight requires

  1. Active YARN applications stay at or below 0.5 on both average and peak for the whole window.
  2. The metric history covers at least 30 days.
  3. The cluster has a known monthly cost.
  4. The cluster has no scheduled deletion configured, such as a max-idle or max-age setting.

Clusters it deliberately ignores

A cluster with scheduled deletion already set is treated as ephemeral and managed, so it is not flagged. Clusters running Presto, Trino, HBase or Flink as optional components are skipped as well, because those services do work outside YARN and would look idle on this metric. When the YARN metric is missing for a cluster, there is no finding.

Saving the whole cluster cost

Terminal window
saving = current monthly cluster cost (VMs + Dataproc fee)
cost after deletion = 0

Deleting the cluster and preventing the next one

  1. Check scheduled jobs, workflow templates and orchestration tools such as Cloud Composer for references to the cluster.
  2. Copy anything left in cluster HDFS to Cloud Storage.
  3. Delete it: gcloud dataproc clusters delete CLUSTER_NAME --region=REGION.
  4. For future clusters, add scheduled deletion at creation, for example gcloud dataproc clusters create CLUSTER_NAME --region=REGION --delete-max-idle=2h, so an idle cluster removes itself.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·