Cloud Run services capped below 10 concurrent requests per instance
What does ZopNight detect here?
Cloud Run runs more instances when each one may handle fewer requests at once. ZopNight flags services whose maximum concurrency is set below 10, measures their peak active `container/instance_count` over 30 days, and prices how many warm instances a concurrency of 80 would remove, reporting only savings of $5 a month or more.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1208 |
| Category | rightsizing |
| Severity | low |
| Metric | container/instance_count (active) |
| Threshold | max concurrency below 10, saving of at least $5 |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | run.services.list · run.services.get · monitoring.timeSeries.list |
Where it applies
Low concurrency multiplies instance count
The concurrency guide explains the setting: the maximum number of requests one instance processes at the same time, up to 1,000. Services deployed with gcloud or Terraform default to 80 times the number of vCPUs; the console default is 80. Lower the limit and Cloud Run needs more instances for the same load. Google’s cost note puts it plainly: “A higher concurrency setting lets fewer instances handle the same request volume, which can reduce costs.”
Low values are sometimes right. Google itself suggests concurrency 1 for code that cannot process parallel requests. They are also often left over from debugging or copied from a template.
Reading concurrency and instance counts yourself
gcloud run services list --region=us-central1gcloud run services describe SERVICE --region=us-central1 --format=yamlcontainerConcurrency in the revision spec is the limit. Then chart
run.googleapis.com/container/instance_count filtered to state="active" in Metrics Explorer to
see how many instances actually served traffic at peak.
Signals ZopNight combines
The service’s max concurrency must be set and below 10. ZopNight takes the peak active instance count from the last 30 days, and requires at least 7 days of peak data before trusting it. It also needs the service’s vCPU and memory settings and per-second rates for both.
Cases where it holds back
No finding appears when concurrency is not set as a limit, when no active instance series exists, when peak coverage is under a week, or when raising concurrency would not remove a single instance. Idle instances are excluded on purpose, because warm minimum instances are priced by the min-instance rules, and counting them here would double the saving. A result under $5 a month is dropped.
Instances a concurrency of 80 would remove
instances at target = ceil(peak active instances x current concurrency / 80)instances saved = peak active instances - instances at targetsaving = instances saved x (vCPU x vCPU rate + GiB x memory rate) x seconds per month, capped at the service's costRaising concurrency without breaking the service
- Check that the code is safe for parallel requests: shared globals, connection pools and per-request temp files are the usual traps.
- Review request latency and CPU per instance; a CPU-bound single-threaded app gains little.
- Raise the limit in steps:
gcloud run services update SERVICE --region=us-central1 --concurrency=40. - Watch latency and error rate after each step before going higher.