Interactive Databricks clusters whose workers run on full-price on-demand VMs
What does ZopNight detect here?
ZopNight flags running all-purpose Databricks clusters whose availability setting starts with `ON_DEMAND`, so every worker VM pays the full cloud rate. The saving covers only the workers, never the driver, and uses the live gap between on-demand and spot rates for that instance type; with no live spot rate, no finding appears.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-2314 · RC-2414 · RC-2214 |
| Category | discount |
| Severity | medium |
| Metric | none — pure configuration read |
| Threshold | availability starts with ON_DEMAND |
| Source | ZopNight |
| Permissions used | GET /api/2.1/clusters/list · GET /api/2.1/clusters/get |
Where it applies
Where the on-demand setting lives on each cloud
Every classic Databricks cluster carries a cloud-specific availability field. The
Clusters API lists the choices:
SPOT, ON_DEMAND and SPOT_WITH_FALLBACK in aws_attributes, with spot-with-fallback as
the AWS default; SPOT_AZURE, ON_DEMAND_AZURE and SPOT_WITH_FALLBACK_AZURE in
azure_attributes; and PREEMPTIBLE_GCP, ON_DEMAND_GCP and PREEMPTIBLE_WITH_FALLBACK_GCP
for Google Cloud. On Azure and GCP the on-demand value is the default, so clusters created
without anyone thinking about capacity type end up paying list price for every worker.
Spot capacity changes only the virtual machine part of the bill. Databricks keeps the driver safe regardless: Azure Databricks documents that the first instance is always on-demand and the rest become spot, and that evicted spot workers are replaced with new spot capacity or, failing that, on-demand instances.
Auditing capacity type across a workspace
List running clusters and look at the availability value in each cloud attribute block:
databricks clusters list --cluster-states RUNNING -o json \ | jq -r '.[] | [.cluster_id, .cluster_name, .cluster_source, (.aws_attributes.availability // .azure_attributes.availability // "see gcp_attributes")] | @tsv'For a single cluster, databricks clusters get 0101-123456-abcde12 -o json shows the full
attribute block, including first_on_demand, the number of initial nodes kept on on-demand.
What ZopNight requires before suggesting spot workers
- The cluster was created interactively, from the UI or API. Job, pipeline, SQL and model serving clusters are managed by Databricks and are left out.
- Its status is not stopped or in error.
- The availability value begins with
ON_DEMAND, which matches the AWS, Azure and GCP spellings. Spot and spot-with-fallback clusters are already converted and are skipped. - There is at least one worker. A fixed cluster with zero workers is skipped, while an autoscaling cluster with a maximum of one or more workers still qualifies, using that maximum as the worker count.
Clusters that get no spot suggestion
With no availability value recorded, the rule has nothing to judge and says nothing. It also stays silent when the cluster has no monthly price, when no live spot rate exists for the worker instance type in that region, or when the computed discount is not positive. There is no default discount percentage to fall back on, by design.
Pricing the worker-only discount
worker share = workers / (workers + 1)saving = cluster monthly VM cost x worker share x (1 - spot rate / on-demand rate)The driver is counted in the denominator and stays on-demand, so a two-worker cluster can move at most two thirds of its VM cost. The spot fraction comes from live tier rates for the instance, not an assumption.
Moving the workers to spot capacity
- Confirm the workload can lose a worker mid-run. Batch Spark stages that retry tolerate this; long stateful sessions may not.
- Edit the cluster and set availability to the spot-with-fallback value for your cloud, so reclaimed workers are replaced with on-demand capacity instead of stalling.
- Keep
first_on_demandat 1 or more so the driver never runs on spot. - Put the setting in a cluster policy so new clusters start on spot-with-fallback.