Skip to main content
compliance · gcp

Dataproc clusters labelled autoscaling=false that hold a fixed worker count

resource types
1
rule IDs covered
1
severity
medium

What does ZopNight detect here?

Dataproc clusters without an autoscaling policy keep the same number of workers whether one job or fifty is queued, so quiet periods pay for peak capacity. ZopNight flags clusters carrying the label `autoscaling=false` and recommends attaching a policy with `gcloud dataproc clusters update --autoscaling-policy`, reported as a compliance finding without a dollar estimate.

Signal and threshold

How ZopNight evaluates Dataproc clusters labelled autoscaling=false that hold a fixed worker count.
Field Value
Rule IDsRC-1211
Categorycompliance
Severitymedium
Metricnone — pure configuration read
Thresholdcluster label autoscaling=false
SourceZopNight
Permissions useddataproc.clusters.list

What a fixed-size Dataproc cluster pays for between jobs

A Dataproc cluster bills for every worker VM it holds, busy or not. An autoscaling policy lets the cluster add and remove workers based on pending and available YARN memory. The Dataproc autoscaling guide describes a policy as a reusable configuration that sets scaling boundaries, frequency and aggressiveness, and it scales worker nodes only, never the master, and only horizontally.

Google recommends autoscaling for clusters that keep their data in Cloud Storage or BigQuery, that run many jobs, or that need to scale up for a single large job. A cluster sized for the nightly peak and left at that size all day is paying for the peak all day.

Finding clusters without a policy

Terminal window
gcloud dataproc clusters list --region=REGION \
--format="table(clusterName, labels.autoscaling, config.autoscalingConfig.policyUri)"

An empty policy column means the cluster has no autoscaling policy attached. The labels column shows the autoscaling label this rule reads.

Why the label is what triggers it

This check is label-driven. ZopNight reads the labels on each Dataproc cluster, and the rule fires only when the cluster carries the label autoscaling with the value false. Many teams use such a label to record a sizing decision; ZopNight treats false as a declared fixed-size cluster that should be reviewed. There is no metric, utilisation threshold or time window in this rule.

Clusters that are not reported

A cluster with no autoscaling label is not flagged, whether or not it has a policy, and neither is one labelled true. Keep that in mind when reading results: the command above is the way to see every cluster with no policy at all. A cluster that sits unused is a different problem, covered by GCP Dataproc Cluster Idle. Google itself says autoscaling is not meant to shrink an idle cluster to minimum size; deleting and recreating ephemeral clusters is the better pattern there.

Why no saving is shown

The finding is categorised as compliance and carries no dollar estimate, because the size of the saving depends on how much your job load varies, which a label cannot tell. The waste is real when the load is spiky: every hour between peaks bills the full worker count.

Attaching an autoscaling policy

  1. Check the workload suits autoscaling. Google does not recommend it for on-cluster HDFS, YARN node labels or Spark Structured Streaming.

  2. Write a policy file. Keep primary workers fixed and let secondary workers flex, for example workerConfig with equal minInstances and maxInstances, and a secondaryWorkerConfig maxInstances.

  3. Import it into the cluster’s region:

    Terminal window
    gcloud dataproc autoscaling-policies import POLICY_NAME \
    --source=policy.yaml --region=REGION
  4. Attach it to the existing cluster (this is not available in the console):

    Terminal window
    gcloud dataproc clusters update CLUSTER_NAME \
    --autoscaling-policy=POLICY_NAME --region=REGION
  5. Update the cluster’s autoscaling label to true so the record matches reality.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·