Interactive Databricks clusters with auto-termination off or set above 120 minutes
What does ZopNight detect here?
ZopNight flags running all-purpose Databricks clusters where `autotermination_minutes` is 0, meaning never, or above 120 minutes. The saving is the cluster's monthly cost times its measured idle share of the week, reduced for long-but-enabled timeouts by 1 minus 120 divided by the setting, and nothing is shown without price and idle data.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-2301 · RC-2401 · RC-2201 |
| Category | schedule |
| Severity | medium |
| Metric | none — pure configuration read |
| Threshold | autotermination_minutes = 0 or > 120 |
| Source | ZopNight |
| Permissions used | GET /api/2.1/clusters/list · GET /api/2.1/clusters/get |
Where it applies
An idle cluster bills both DBUs and VMs until it stops
Databricks is explicit that idle compute is not free. The
compute management guide says idle
compute continues to accumulate DBU and cloud instance charges during the inactivity period before
termination. Through the Clusters API, a cluster created without the field never terminates on its own. Valid inactivity periods run from
10 to 10,000 minutes, roughly a week, and 0 switches the feature off.
A cluster counts as inactive once all commands have finished, including Spark jobs, Structured Streaming, JDBC calls and web terminal activity. Commands run over SSH outside the web terminal do not count as activity. So a cluster left with auto-termination off keeps billing overnight and over weekends after the last notebook cell finishes.
Checking the setting across clusters
databricks clusters list -o json \ | jq -r '.[] | [.cluster_id, .cluster_name, .cluster_source, .autotermination_minutes] | @tsv' \ | sort -t$'\t' -k4 -nRows with 0 never stop on their own; rows above 120 wait more than two hours. The same field
appears in databricks clusters get CLUSTER_ID -o json if you want the whole configuration.
Settings that trigger the finding
The rule reads the configured timeout on interactive clusters (created from the UI or API) that are not stopped or in error. A value of 0 counts as disabled and anything above 120 minutes counts as too loose. Exactly 120 is treated as healthy. Clusters created by jobs, pipelines, SQL warehouses and model serving are excluded, since the job, pipeline, warehouse or endpoint that created them controls their lifecycle.
When the timeout is not reported
A cluster with no recorded timeout is not judged. The finding also needs a dollar figure: ZopNight must have a monthly price for the cluster and a measured idle share of its week. Without either, or if the calculation comes out at zero, nothing is shown rather than a $0 row.
From idle share to monthly saving
disabled: saving = monthly cost x idle share of the weektoo long: saving = monthly cost x idle share x (1 - 120 / configured minutes)A 240-minute timeout already reclaims part of the idle time, so only half the idle share counts (1 - 120/240). When this cluster also has a fixed-size finding from Cluster Autoscaling Disabled, the combined savings are capped at the cluster’s monthly cost.
Setting a sensible timeout
- Pick an inactivity period between 10 and 120 minutes; shorter for personal exploration, longer for clusters behind a dashboard.
- Edit the cluster in Compute and enter the period in the auto termination field.
- Make it the default with a cluster policy that fixes or ranges
autotermination_minutes, so new clusters cannot opt out.