Autoscaling Databricks clusters with a minimum of 4+ workers that often idle below 20% CPU
What does ZopNight detect here?
ZopNight looks for running all-purpose Databricks clusters with autoscaling on and a minimum of 4 or more workers, where at least 10% of 30 days of `CPUUtilization` readings are under 20%. That floor bills around the clock, so ZopNight prices lowering it to 2, halved because the cluster only sometimes sits at its minimum.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-2306 · RC-2406 · RC-2206 |
| Category | rightsizing |
| Severity | low |
| Metric | CPUUtilization |
| Threshold | min workers >= 4 and 10% of readings below 20% CPU |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | GET /api/2.1/clusters/list · GET /api/2.1/clusters/get |
Where it applies
The autoscale minimum is a fixed cost in disguise
Autoscaling only moves a cluster between its bounds. Whatever is typed into Min stays up for as long as the cluster runs, busy or not, so a minimum of 8 workers is a fixed 8-worker cluster during every quiet hour. Databricks even works to protect that floor: its compute reference says that when the cloud provider terminates instances and the cluster drops below its minimum, Databricks keeps retrying to provision instances to get back to it.
Teams often set a high minimum once, to make the first query of the day fast, and never revisit it.
Comparing autoscale bounds with real load
databricks clusters list --cluster-states RUNNING -o json \ | jq -r '.[] | select(.autoscale != null) | [.cluster_id, .cluster_name, .autoscale.min_workers, .autoscale.max_workers] | @tsv'For any cluster with a minimum of 4 or more, check the CPU chart on its Metrics tab. Long flat stretches well under 20% mean the floor is doing nothing.
Gates on the autoscale floor
- Interactive cluster (UI or API), not stopped and not in error.
- Both autoscale bounds recorded, with the maximum above zero. Clusters with autoscaling off are handled by Cluster Autoscaling Disabled.
- Minimum of at least 4 workers, which leaves room above the 2-worker target.
- A cluster CPU series (the average of
CPUUtilizationover its worker VMs) with usable readings in the last 30 days, of which at least 10% are below 20%. - A monthly price and a known worker count.
When the floor is left as it is
Busy clusters, where fewer than 10% of readings fall under 20%, are left alone: the floor is being used. A missing CPU series means no finding rather than a guess. The series exists for AWS and Azure worker VMs only, so this rule does not fire for Databricks on Google Cloud. On AWS, a cluster whose workers come from an instance pool has no series either, since pool-launched EC2 instances do not carry the cluster’s tags.
Sizing the saving from a lower minimum
node rate = cluster monthly cost / (workers + 1)nodes reclaimed = (min workers - 2) / 2 x share of readings below 20% CPUsaving = node rate x nodes reclaimedHalving the reduction is deliberate. An autoscaling cluster spends only part of its time at the floor, so counting the full drop from, say, 8 to 2 every hour would overstate the result.
Trimming the floor
- Review 14 days of CPU for the cluster and note how many workers the quiet periods need.
- Edit the cluster and lower Min toward 2, leaving Max alone for peaks.
- If CPU is low even at peak, try a smaller worker node type as well.
- Recheck after a week: if the cluster now hovers at 2 workers with healthy CPU, the old floor was pure overhead.