Databricks model serving endpoints with scale-to-zero turned off
What does ZopNight detect here?
ZopNight examines Databricks model serving endpoints backed by provisioned compute where `scale_to_zero_enabled` is false, so capacity stays up between requests. A finding needs a real monthly dollar figure for the idle hours, and ZopNight does not yet measure the endpoint's serving cost or traffic, so it shows nothing today rather than a $0 advisory.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-2304 · RC-2404 · RC-2204 |
| Category | idle |
| Severity | low |
| Metric | none — pure configuration read |
| Threshold | scale_to_zero_enabled = false |
| Source | ZopNight |
| Permissions used | GET /api/2.0/serving-endpoints · GET /api/2.0/serving-endpoints/{name} |
Where it applies
Provisioned serving capacity that never scales down
A custom model endpoint in Databricks Model Serving runs on compute sized by workload_size.
The serving endpoint API
defines Small as 4 provisioned concurrency, Medium as 8 to 16 and Large as 16 to 64. The
scale_to_zero_enabled flag decides what happens when no requests arrive: with it on, the endpoint
releases its compute while idle; with it off, the provisioned capacity stays up around the clock.
Databricks is candid about the trade-off. Its custom endpoint guide says an endpoint that has scaled to zero pays a cold start on the next request, and that scale to zero is not recommended for production endpoints because capacity is not guaranteed while scaled down. For development, batch-style or occasional traffic, that warning rarely applies.
Reading scale-to-zero on every endpoint
databricks serving-endpoints list -o json | jq -r '.[].name' | while read -r ep; do databricks serving-endpoints get "$ep" -o json \ | jq -r --arg ep "$ep" '.config.served_entities[]? | [$ep, .name, .workload_size, (.scale_to_zero_enabled|tostring)] | @tsv'doneRows ending in false are endpoints that keep their capacity when traffic stops. Check each
one’s request pattern in the endpoint’s metrics before deciding.
Which endpoints are in scope
Only endpoints running on provisioned compute are considered. Pay-per-token Foundation Model API endpoints and external model endpoints have no dedicated compute to release, so they are excluded even though they also report scale-to-zero as off. Endpoints that are stopped or in error are skipped, and an endpoint whose scale-to-zero value was not collected is not judged.
Why you may see no finding for an endpoint
Every idle-cost finding from ZopNight must carry a real dollar figure. For a serving endpoint that figure depends on the serving DBU rate for its workload size and on how many hours a month it actually handled requests. ZopNight does not yet collect either input, so today the rule stays silent on every endpoint instead of posting a zero-dollar warning. The absence of a finding here does not mean scale-to-zero is set.
How the idle cost would be sized
idle cost = serving DBU rate x DBUs per hour for the workload size x (730 - active hours)The 730 is the hours in an average month; only the hours with no traffic count as recoverable.
Enabling scale-to-zero safely
- Open Serving, pick the endpoint and review its request volume by hour.
- If traffic is intermittent and a cold start is acceptable, edit the served entity and turn on
scale to zero, or send
"scale_to_zero_enabled": truefor it withdatabricks serving-endpoints update-config <name> --json @config.json. - Keep scale-to-zero off for latency-critical endpoints with steady traffic, and right-size their workload size instead.