Cloud Run services paying to keep minimum instances warm with no traffic
What does ZopNight detect here?
Cloud Run services are flagged when they pin one or more minimum instances yet `run.googleapis.com/request_count` shows effectively no traffic for 30 days. Google bills idle minimum instances for CPU and memory at the idle rate, so ZopNight prices the warm capacity from the service's vCPU and memory size and recommends scaling the minimum to 0.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-149 |
| Category | idle |
| Severity | medium |
| Metric | run.googleapis.com/request_count |
| Threshold | min instances 1 or more with request rate under 0.001 |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | run.services.list · run.services.get · monitoring.timeSeries.list |
Where it applies
Warm instances that nobody calls
Cloud Run normally scales to zero: an instance not handling requests becomes idle and is shut down. Setting a minimum keeps instances running anyway. Google’s minimum instances page notes that instances kept running this way incur billing costs, and the pricing page lists a separate idle rate per vCPU-second and GiB-second for them, adding that idle instances which are not minimum instances are not charged.
Minimums make sense for latency-sensitive services with steady traffic. On a service with no requests for a month, they are a standing charge for readiness nobody uses.
Checking a service’s minimum and its traffic
The minimum appears as a minScale annotation in the service YAML:
gcloud run services describe SERVICE --region=REGION --format=yaml | grep -i minscaleFor traffic, chart run.googleapis.com/request_count for the service over 30 days in Metrics
Explorer. Note that this metric counts only requests that reach the container, so requests
rejected by IAM do not appear.
Conditions ZopNight checks
- The service’s minimum instance count is 1 or more.
- Request activity over the 30-day window averages below 0.001, effectively none.
- The service’s vCPU and memory size are known, along with live per-second Cloud Run rates for the region.
- The service has a known monthly cost.
When it declines to flag
With no request data at all, the rule stays silent rather than assuming no traffic. It also stays silent when the minimum is 0, when CPU or memory size cannot be read, or when either per-second rate is missing, because the idle charge cannot then be priced. Dev and test services that keep a minimum regardless of traffic are a separate check: Cloud Run Dev/Test Service Has Min Instances Above 0.
Pricing the idle warm capacity
saving = min instances x (vCPU x CPU rate per second + memory GiB x memory rate per second) x 2,592,000 seconds per monthsaving is capped at the service's current monthly costThe 2,592,000 seconds is a 30-day month. Request-driven capacity is unchanged, because Cloud Run still scales up when traffic arrives.
Scaling the minimum to zero
- Confirm in the metric above that the service really sees no traffic.
- Set the service-level minimum to zero without a new revision:
gcloud run services update SERVICE --region=REGION --min=0. - If the minimum was set per revision, use
--min-instances=0instead, which deploys a new revision. - Watch cold-start latency on the next real request and decide whether that is acceptable.