Skip to main content
governance · databricks

Interactive Databricks clusters created freehand with no cluster policy attached

resource types
1
rule IDs covered
3
severity
low

What does ZopNight detect here?

ZopNight reports running all-purpose Databricks clusters whose `policy_id` was read and found empty, meaning nobody bounded node type, worker count, tags or auto-termination when the cluster was created. The finding is governance only and carries no dollar saving; ZopNight skips any cluster where the policy field was simply not collected.

Signal and threshold

How ZopNight evaluates Interactive Databricks clusters created freehand with no cluster policy attached.
Field Value
Rule IDsRC-2319 · RC-2419 · RC-2219
Categorygovernance
Severitylow
Metricnone — pure configuration read
Thresholdpolicy_id empty
SourceZopNight
Permissions usedGET /api/2.1/clusters/list · GET /api/2.0/policies/clusters/list

What a missing policy leaves uncontrolled

A cluster policy is the guardrail Databricks gives admins for compute creation. Databricks describes policies as a way to limit users to prescribed settings, limit how many clusters each user can create, and control cost by limiting per-cluster maximum cost. Users who hold the Unrestricted cluster creation entitlement get an Unrestricted policy that lets them configure anything, and that is where freehand clusters come from.

Without a policy nothing stops a cluster from being built on the largest node type, with a high fixed worker count, no cost tags and no idle shutdown. None of that costs money on the day it is created; it costs money every hour afterwards.

Finding policy-less clusters from the CLI

The cluster record exposes policy_id. List clusters where it is absent or blank:

Terminal window
databricks clusters list -o json \
| jq -r '.[] | select((.policy_id // "") == "") | [.cluster_id, .cluster_name, .cluster_source] | @tsv'

Then see which policies already exist and could be attached:

Terminal window
databricks cluster-policies list -o json | jq -r '.[] | [.policy_id, .name] | @tsv'

The reverse filter, databricks clusters list --policy-id <id>, is handy for checking which clusters a policy already governs.

The evidence this rule insists on

ZopNight fires only when it actually read the cluster’s policy field and the value was empty. The cluster must also be interactive (created from the UI or API, not by a job, pipeline, SQL warehouse or model serving) and must not be stopped or in error, so the finding always concerns compute that is live.

Why a blank field is not always a finding

A missing field and an empty field are different things. If the policy field was never collected for a cluster, or ZopNight holds no configuration for it at all, the rule does not raise a finding, because it cannot tell “no policy” apart from “unknown”. Terminated clusters are skipped too; the gap matters again when someone restarts them.

No saving is attached, only exposure

This is a governance finding, so the saving shown is zero. The value is preventive: a policy caps the settings that other rules measure after the fact, such as Cluster Missing/Weak Auto-Termination and Missing Cost-Allocation Tags.

Putting clusters under a policy

  1. Start from a policy family or write a definition. Policy attributes can be fixed, range or allowlist, for example a fixed autotermination_minutes, a range on autoscale.max_workers, an allowlist on node_type_id, and fixed custom_tags.team; see the policy definition reference.
  2. Assign it to existing all-purpose compute with Add compute to policy, which previews the configuration changes before applying them.
  3. Grant users the policy and remove Unrestricted cluster creation from those who do not need it, so new clusters cannot be created freehand.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·