Fixed-size Databricks clusters with 2+ workers that sit below 20% CPU often enough to shrink
What does ZopNight detect here?
ZopNight targets running all-purpose Databricks clusters with autoscaling off and at least 2 fixed workers, where 10% or more of the last 30 days of `CPUUtilization` readings fall under 20%. The saving is the cost of every worker above one, scaled by that low-load share, because an autoscaling cluster could have shed them.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-2308 · RC-2408 · RC-2208 |
| Category | rightsizing |
| Severity | low |
| Metric | CPUUtilization |
| Threshold | 10% of readings below 20% CPU |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | GET /api/2.1/clusters/list · GET /api/2.1/clusters/get |
Where it applies
A fixed worker count bills for peak all day
When autoscaling is off, a Databricks cluster runs the exact number of workers typed into the Workers field for as long as it is up, whether a notebook is crunching a join or nobody is attached. Databricks notes that autoscaling can reduce overall cost compared to a statically sized cluster, because workers are added for demanding phases and removed when no longer needed. On Premium workspaces, optimized autoscaling scales an all-purpose cluster down after 150 seconds of underutilization.
Finding static clusters in a workspace
A cluster without an autoscale block uses num_workers as a fixed size:
databricks clusters list --cluster-states RUNNING -o json \ | jq -r '.[] | select(.autoscale == null) | [.cluster_id, .cluster_name, .num_workers, .node_type_id] | @tsv'To judge whether one would benefit, open Compute, choose the cluster and review the Metrics tab for CPU across a normal working week.
Evidence required before a finding
- The cluster is interactive, meaning created from the UI or API, and is not stopped or in error.
- Autoscaling is recorded as off, with no maximum worker count set.
- It has at least 2 workers. A single-worker cluster has nothing meaningful to shed.
- ZopNight holds a CPU series for the cluster, built by averaging
CPUUtilizationacross its worker VMs, with usable datapoints in the last 30 days. - At least 10% of those datapoints are below 20% CPU.
- The cluster has a monthly price.
Situations where no finding is produced
If the autoscaling setting was not collected, ZopNight does not assume the cluster is static. Steady clusters, where fewer than 10% of readings dip below 20%, get no finding because there is no defensible scale-down. Per-cluster CPU is built from the worker VMs on AWS and Azure; on Google Cloud there is no such series, so the rule stays silent there. On AWS, clusters launched from a pool are also silent, because their EC2 instances carry pool tags rather than cluster tags and cannot be matched to the cluster. Clusters that already autoscale belong to Databricks Cluster Oversized instead.
Pricing the workers autoscaling would drop
node rate = cluster monthly cost / (workers + 1)shed nodes = workers - 1saving = node rate x shed nodes x share of readings below 20% CPUThe driver and one worker are assumed to keep running. If the same cluster also has a loose auto-termination finding, ZopNight caps the combined savings at the cluster’s monthly cost.
Switching the cluster to autoscaling
- Edit the cluster and tick Enable autoscaling.
- Set Min near the quiet-hour need, often 1 or 2, and Max at the old fixed size. Databricks resizes into those bounds immediately; a 12-worker cluster set to 5 to 10 drops to 10.
- Watch the cluster metrics for a week and adjust the bounds.