Scheduled Databricks jobs running on interactive all-purpose clusters
What does ZopNight detect here?
ZopNight finds Databricks jobs whose tasks point at an existing all-purpose cluster through `existing_cluster_id`, the interactive tier Databricks prices above jobs compute. The saving would be the job's measured DBU-hours over 30 days multiplied by the regional price gap between all-purpose and jobs DBUs. Neither number is collected yet, so no finding is shown today.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-2313 · RC-2413 · RC-2213 |
| Category | rightsizing |
| Severity | high |
| Metric | DBU-hours per job |
| Threshold | targets an all-purpose cluster |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | GET /api/2.2/jobs/list · GET /api/2.2/jobs/runs/list |
Where it applies
Interactive compute is the expensive way to run a schedule
Databricks prices compute by workload type. Its cost optimization guide says non-interactive workloads cost significantly less on job compute than on all-purpose compute. The DBUs consumed are the same work either way; only the rate changes. On top of that, the jobs compute guide states plainly that Databricks does not recommend running production jobs on all-purpose compute.
A job cluster exists only for its run and terminates when the job completes, while an all-purpose cluster that a job targets has to be kept around for the schedule, often with a long idle timeout.
Listing jobs that target existing clusters
A task that names existing_cluster_id runs on an all-purpose cluster; new_cluster and
job_cluster_key use jobs compute:
databricks jobs list --expand-tasks -o json \ | jq -r '.[] | .job_id as $id | .settings.name as $n | .settings.tasks[]? | select(.existing_cluster_id) | [$id, $n, .task_key, .existing_cluster_id] | @tsv'Check how often each one runs before prioritizing:
databricks jobs list-runs --job-id 123456789 --limit 25 -o json | jq 'length'Inputs ZopNight needs to price the move
- The job is recorded as targeting all-purpose compute.
- Its DBU-hours from runs over the last 30 days are measured and above zero.
- A price gap per DBU-hour between all-purpose and jobs compute is known for the job’s region and is positive.
Jobs that get no recommendation
A job that ZopNight knows ran zero times in 30 days is skipped, because moving a job that never runs saves nothing; see Orphaned Job for those. Jobs already on job clusters or serverless are fine. ZopNight does not yet collect the measured DBU-hours or the regional rate gap, so today the rule stays silent on every job instead of estimating, and an all-purpose job will not show a finding. The CLI check above is the complete list.
How the saving is calculated
saving per month = DBU-hours used by the job in 30 days x (all-purpose $ per DBU-hour - jobs $ per DBU-hour)Only the DBU rate changes in the move. Cloud VM cost for the run is not counted, although dropping a pinned all-purpose cluster the job no longer needs can save more on top.
Moving the job to jobs compute
- Copy the all-purpose cluster’s runtime, node type and libraries into a
new_clusterdefinition, or a sharedjob_clustersentry if several tasks run in sequence. - Change each task from
existing_cluster_idto that job cluster and run it once by hand. - For notebook and Python tasks, consider serverless jobs, which Databricks recommends for those task types.
- Keep the all-purpose cluster for interactive notebook work, or terminate it if the job was its only user.