Skip to main content
rightsizing · databricks

Fixed-size Databricks clusters with 2+ workers that sit below 20% CPU often enough to shrink

resource types
1
rule IDs covered
3
severity
low

What does ZopNight detect here?

ZopNight targets running all-purpose Databricks clusters with autoscaling off and at least 2 fixed workers, where 10% or more of the last 30 days of `CPUUtilization` readings fall under 20%. The saving is the cost of every worker above one, scaled by that low-load share, because an autoscaling cluster could have shed them.

Signal and threshold

How ZopNight evaluates Fixed-size Databricks clusters with 2+ workers that sit below 20% CPU often enough to shrink.
Field Value
Rule IDsRC-2308 · RC-2408 · RC-2208
Categoryrightsizing
Severitylow
MetricCPUUtilization
Threshold10% of readings below 20% CPU
Evaluation window30d
SourceZopNight
Permissions usedGET /api/2.1/clusters/list · GET /api/2.1/clusters/get

A fixed worker count bills for peak all day

When autoscaling is off, a Databricks cluster runs the exact number of workers typed into the Workers field for as long as it is up, whether a notebook is crunching a join or nobody is attached. Databricks notes that autoscaling can reduce overall cost compared to a statically sized cluster, because workers are added for demanding phases and removed when no longer needed. On Premium workspaces, optimized autoscaling scales an all-purpose cluster down after 150 seconds of underutilization.

Finding static clusters in a workspace

A cluster without an autoscale block uses num_workers as a fixed size:

Terminal window
databricks clusters list --cluster-states RUNNING -o json \
| jq -r '.[] | select(.autoscale == null) | [.cluster_id, .cluster_name, .num_workers, .node_type_id] | @tsv'

To judge whether one would benefit, open Compute, choose the cluster and review the Metrics tab for CPU across a normal working week.

Evidence required before a finding

  1. The cluster is interactive, meaning created from the UI or API, and is not stopped or in error.
  2. Autoscaling is recorded as off, with no maximum worker count set.
  3. It has at least 2 workers. A single-worker cluster has nothing meaningful to shed.
  4. ZopNight holds a CPU series for the cluster, built by averaging CPUUtilization across its worker VMs, with usable datapoints in the last 30 days.
  5. At least 10% of those datapoints are below 20% CPU.
  6. The cluster has a monthly price.

Situations where no finding is produced

If the autoscaling setting was not collected, ZopNight does not assume the cluster is static. Steady clusters, where fewer than 10% of readings dip below 20%, get no finding because there is no defensible scale-down. Per-cluster CPU is built from the worker VMs on AWS and Azure; on Google Cloud there is no such series, so the rule stays silent there. On AWS, clusters launched from a pool are also silent, because their EC2 instances carry pool tags rather than cluster tags and cannot be matched to the cluster. Clusters that already autoscale belong to Databricks Cluster Oversized instead.

Pricing the workers autoscaling would drop

Terminal window
node rate = cluster monthly cost / (workers + 1)
shed nodes = workers - 1
saving = node rate x shed nodes x share of readings below 20% CPU

The driver and one worker are assumed to keep running. If the same cluster also has a loose auto-termination finding, ZopNight caps the combined savings at the cluster’s monthly cost.

Switching the cluster to autoscaling

  1. Edit the cluster and tick Enable autoscaling.
  2. Set Min near the quiet-hour need, often 1 or 2, and Max at the old fixed size. Databricks resizes into those bounds immediately; a 12-worker cluster set to 5 to 10 drops to 10.
  3. Watch the cluster metrics for a week and adjust the bounds.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·