Skip to main content
schedule · databricks

Interactive Databricks clusters with auto-termination off or set above 120 minutes

resource types
1
rule IDs covered
3
severity
medium

What does ZopNight detect here?

ZopNight flags running all-purpose Databricks clusters where `autotermination_minutes` is 0, meaning never, or above 120 minutes. The saving is the cluster's monthly cost times its measured idle share of the week, reduced for long-but-enabled timeouts by 1 minus 120 divided by the setting, and nothing is shown without price and idle data.

Signal and threshold

How ZopNight evaluates Interactive Databricks clusters with auto-termination off or set above 120 minutes.
Field Value
Rule IDsRC-2301 · RC-2401 · RC-2201
Categoryschedule
Severitymedium
Metricnone — pure configuration read
Thresholdautotermination_minutes = 0 or > 120
SourceZopNight
Permissions usedGET /api/2.1/clusters/list · GET /api/2.1/clusters/get

An idle cluster bills both DBUs and VMs until it stops

Databricks is explicit that idle compute is not free. The compute management guide says idle compute continues to accumulate DBU and cloud instance charges during the inactivity period before termination. Through the Clusters API, a cluster created without the field never terminates on its own. Valid inactivity periods run from 10 to 10,000 minutes, roughly a week, and 0 switches the feature off.

A cluster counts as inactive once all commands have finished, including Spark jobs, Structured Streaming, JDBC calls and web terminal activity. Commands run over SSH outside the web terminal do not count as activity. So a cluster left with auto-termination off keeps billing overnight and over weekends after the last notebook cell finishes.

Checking the setting across clusters

Terminal window
databricks clusters list -o json \
| jq -r '.[] | [.cluster_id, .cluster_name, .cluster_source, .autotermination_minutes] | @tsv' \
| sort -t$'\t' -k4 -n

Rows with 0 never stop on their own; rows above 120 wait more than two hours. The same field appears in databricks clusters get CLUSTER_ID -o json if you want the whole configuration.

Settings that trigger the finding

The rule reads the configured timeout on interactive clusters (created from the UI or API) that are not stopped or in error. A value of 0 counts as disabled and anything above 120 minutes counts as too loose. Exactly 120 is treated as healthy. Clusters created by jobs, pipelines, SQL warehouses and model serving are excluded, since the job, pipeline, warehouse or endpoint that created them controls their lifecycle.

When the timeout is not reported

A cluster with no recorded timeout is not judged. The finding also needs a dollar figure: ZopNight must have a monthly price for the cluster and a measured idle share of its week. Without either, or if the calculation comes out at zero, nothing is shown rather than a $0 row.

From idle share to monthly saving

Terminal window
disabled: saving = monthly cost x idle share of the week
too long: saving = monthly cost x idle share x (1 - 120 / configured minutes)

A 240-minute timeout already reclaims part of the idle time, so only half the idle share counts (1 - 120/240). When this cluster also has a fixed-size finding from Cluster Autoscaling Disabled, the combined savings are capped at the cluster’s monthly cost.

Setting a sensible timeout

  1. Pick an inactivity period between 10 and 120 minutes; shorter for personal exploration, longer for clusters behind a dashboard.
  2. Edit the cluster in Compute and enter the period in the auto termination field.
  3. Make it the default with a cluster policy that fixes or ranges autotermination_minutes, so new clusters cannot opt out.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·