Skip to main content
advisory · azure

Idle Azure Databricks pool check, retired so pool waste is priced on the pool itself

rule IDs covered
1
severity
medium

What does ZopNight detect here?

ZopNight's Idle Azure Databricks Pool check is retired and raises no findings on any cluster. Azure Databricks charges no DBUs while pool instances sit idle, yet the underlying VMs keep billing, so ZopNight measures that cost where it lives: on each pool's `min_idle_instances` floor and on clusters with weak auto-termination.

Signal and threshold

How ZopNight evaluates Idle Azure Databricks pool check, retired so pool waste is priced on the pool itself.
Field Value
Rule IDsRC-220
Categoryadvisory
Severitymedium
Metricnone — pure configuration read
SourceZopNight
Permissions usedMicrosoft.Databricks/workspaces/read

Idle pool instances skip DBUs but not the VM bill

An Azure Databricks pool keeps a set of ready instances so clusters start and autoscale faster. Microsoft is explicit about how those instances are billed: Azure Databricks does not charge DBUs while instances are idle in the pool, but instance provider billing does apply. In other words, the Databricks meter stops and the Azure virtual machine meter keeps running.

The instances that never go away are the ones counted by Minimum Idle Instances. The pool configuration reference says these instances do not terminate, regardless of the auto termination setting, and that Databricks provisions replacements whenever a cluster takes one. A floor of four idle VMs is therefore four VMs billed around the clock.

Reading pool floors with the Databricks CLI

List the pools in a workspace together with their usage statistics, then inspect a single pool to see its configured floor and idle timeout:

Terminal window
databricks instance-pools list
databricks instance-pools get 0101-120000-pool1234

Look at min_idle_instances and idle_instance_autotermination_minutes in the output. A non-zero floor on a pool that clusters rarely draw from is the waste worth chasing.

Why the cluster-level idle check was switched off

This check was written against Databricks clusters, but the evidence it needed (job runs and running-cluster counts) is recorded per workspace. A workspace-wide series can never be matched to one cluster, so the check could never prove that any particular cluster was idle. It also looked in the wrong place: the money recoverable from warm capacity sits on the pool’s idle floor, not on a cluster.

Rather than emit placeholder findings with a $0 saving, ZopNight retired the check in place.

What makes it fire today: nothing

Every Databricks cluster is still passed to the check, and every evaluation returns no finding. There is no metric, threshold or lookback window to satisfy. The page stays in the catalog so existing links resolve and so it is clear where the replacement coverage lives.

Where ZopNight reports this saving now

Two active rules price the same waste with real dollar figures. Instance Pool Min-Idle Underutilized looks at pools whose warm floor exceeds what clusters actually use, and Cluster Missing/Weak Auto-Termination catches clusters that hold their instances long after work stops. Findings from either rule appear in place of anything this retired check would have produced.

Trimming idle pool spend by hand

  1. Run databricks instance-pools list in each workspace and note pools with a non-zero minimum idle count.
  2. For pools where a slower cluster start is acceptable, set the floor to zero with databricks instance-pools edit, which takes the pool ID, pool name and node type as arguments; pass --min-idle-instances 0, as the pool best practices recommend.
  3. Set Idle Instance Auto Termination long enough to cover gaps between scheduled jobs and no longer.
  4. If jobs need warm instances at a known time, schedule a short starter job just before them instead of holding a permanent floor.
  5. Tag pools so their VM cost can be charged back; pool tags propagate to Azure billing.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·