Databricks instance pools holding warm idle VMs while no cluster uses them
What does ZopNight detect here?
ZopNight flags a Databricks instance pool whose `min_idle_instances` is above 0 when its live stats show `used_count` of 0 and an `idle_count` above 0, meaning warm VMs are running with no cluster attached. The saving is the pool's full monthly cost for that minimum-idle floor, and unpriced pools get no finding.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-2305 · RC-2405 · RC-2205 |
| Category | rightsizing |
| Severity | medium |
| Metric | pool idle_count and used_count |
| Threshold | min_idle_instances > 0, used_count = 0, idle_count > 0 |
| Source | ZopNight |
| Permissions used | GET /api/2.0/instance-pools/list |
Where it applies
Warm pool instances are running VMs
An instance pool keeps ready-to-use cloud instances so clusters start faster. The
min_idle_instances setting is the number of those instances the pool keeps up even when nothing
needs them. The pool best practices guide
is direct about the cost: set Min Idle to 0 to avoid paying for running instances that are not
doing work, at the price of a slower start when a cluster needs a fresh instance.
A warm floor makes sense when clusters draw from the pool all day. When no cluster is using the pool, the floor is a set of virtual machines billed by your cloud provider for nothing.
Reading live pool statistics
The pool API returns a stats object: idle_count is active instances not part of any cluster,
used_count is active instances that are part of one.
databricks instance-pools list -o json \ | jq -r '.[] | [.instance_pool_id, .instance_pool_name, .node_type_id, .min_idle_instances, .stats.idle_count, .stats.used_count, .idle_instance_autotermination_minutes] | @tsv'A pool with a positive minimum, zero used and some idle is carrying pure standing cost at that moment. Check it at several times of day before acting.
Three facts the pool must show
- A minimum idle count above 0. Pools with no floor have nothing to reduce.
- No instances currently in use by a cluster, and at least one warm idle instance right now.
- A monthly price for the pool, which ZopNight calculates from the node rate, the minimum idle count and the hours in the month.
Pools that are not reported
If no cluster is attached but the pool also has no idle instances, there is no current waste. If the pool statistics were not collected, ZopNight treats the counts as zero and does not fire, rather than falling back to judging the configured floor alone. A pool that cannot be priced produces no finding; ZopNight will not show a $0 cost recommendation.
The saving equals the cost of the floor
saving = node rate x min idle instances x hours in monthcost after fix = 0 for the floorThe figure is based on the configured minimum, not on the idle count seen at one moment, because the minimum is what keeps billing month after month. The finding text shows both numbers so the basis is visible.
Lowering the floor without breaking start times
- Look at when clusters actually attach to the pool, for example around scheduled jobs.
- Reduce Min Idle toward 0:
databricks instance-pools edit <pool-id> <pool-name> <node-type-id> --min-idle-instances 0. - Set
--idle-instance-autotermination-minutesso instances above the minimum are released after a short idle period instead of lingering. - If a job needs warm starts, keep a small floor only for the hours it runs, rather than all day.