Skip to main content
discount · databricks

Interactive Databricks clusters whose workers run on full-price on-demand VMs

resource types
1
rule IDs covered
3
severity
medium

What does ZopNight detect here?

ZopNight flags running all-purpose Databricks clusters whose availability setting starts with `ON_DEMAND`, so every worker VM pays the full cloud rate. The saving covers only the workers, never the driver, and uses the live gap between on-demand and spot rates for that instance type; with no live spot rate, no finding appears.

Signal and threshold

How ZopNight evaluates Interactive Databricks clusters whose workers run on full-price on-demand VMs.
Field Value
Rule IDsRC-2314 · RC-2414 · RC-2214
Categorydiscount
Severitymedium
Metricnone — pure configuration read
Thresholdavailability starts with ON_DEMAND
SourceZopNight
Permissions usedGET /api/2.1/clusters/list · GET /api/2.1/clusters/get

Where the on-demand setting lives on each cloud

Every classic Databricks cluster carries a cloud-specific availability field. The Clusters API lists the choices: SPOT, ON_DEMAND and SPOT_WITH_FALLBACK in aws_attributes, with spot-with-fallback as the AWS default; SPOT_AZURE, ON_DEMAND_AZURE and SPOT_WITH_FALLBACK_AZURE in azure_attributes; and PREEMPTIBLE_GCP, ON_DEMAND_GCP and PREEMPTIBLE_WITH_FALLBACK_GCP for Google Cloud. On Azure and GCP the on-demand value is the default, so clusters created without anyone thinking about capacity type end up paying list price for every worker.

Spot capacity changes only the virtual machine part of the bill. Databricks keeps the driver safe regardless: Azure Databricks documents that the first instance is always on-demand and the rest become spot, and that evicted spot workers are replaced with new spot capacity or, failing that, on-demand instances.

Auditing capacity type across a workspace

List running clusters and look at the availability value in each cloud attribute block:

Terminal window
databricks clusters list --cluster-states RUNNING -o json \
| jq -r '.[] | [.cluster_id, .cluster_name, .cluster_source,
(.aws_attributes.availability // .azure_attributes.availability // "see gcp_attributes")] | @tsv'

For a single cluster, databricks clusters get 0101-123456-abcde12 -o json shows the full attribute block, including first_on_demand, the number of initial nodes kept on on-demand.

What ZopNight requires before suggesting spot workers

  • The cluster was created interactively, from the UI or API. Job, pipeline, SQL and model serving clusters are managed by Databricks and are left out.
  • Its status is not stopped or in error.
  • The availability value begins with ON_DEMAND, which matches the AWS, Azure and GCP spellings. Spot and spot-with-fallback clusters are already converted and are skipped.
  • There is at least one worker. A fixed cluster with zero workers is skipped, while an autoscaling cluster with a maximum of one or more workers still qualifies, using that maximum as the worker count.

Clusters that get no spot suggestion

With no availability value recorded, the rule has nothing to judge and says nothing. It also stays silent when the cluster has no monthly price, when no live spot rate exists for the worker instance type in that region, or when the computed discount is not positive. There is no default discount percentage to fall back on, by design.

Pricing the worker-only discount

Terminal window
worker share = workers / (workers + 1)
saving = cluster monthly VM cost x worker share x (1 - spot rate / on-demand rate)

The driver is counted in the denominator and stays on-demand, so a two-worker cluster can move at most two thirds of its VM cost. The spot fraction comes from live tier rates for the instance, not an assumption.

Moving the workers to spot capacity

  1. Confirm the workload can lose a worker mid-run. Batch Spark stages that retry tolerate this; long stateful sessions may not.
  2. Edit the cluster and set availability to the spot-with-fallback value for your cloud, so reclaimed workers are replaced with on-demand capacity instead of stalling.
  3. Keep first_on_demand at 1 or more so the driver never runs on spot.
  4. Put the setting in a cluster policy so new clusters start on spot-with-fallback.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·