Skip to main content
compliance · gcp

Production Cloud SQL instances running in a single zone with no standby for failover

resource types
1
rule IDs covered
1
severity
high

What does ZopNight detect here?

ZopNight flags Cloud SQL instances that look like production (a name containing prod, production or live, or a production env label) but whose `availabilityType` is not `REGIONAL`. A zonal instance has no standby in a second zone, so a zone failure takes the database down until the zone recovers or someone restores from backup.

Signal and threshold

How ZopNight evaluates Production Cloud SQL instances running in a single zone with no standby for failover.
Field Value
Rule IDsRC-1240
Categorycompliance
Severityhigh
Metricnone — pure configuration read
Thresholdproduction instance with availabilityType other than REGIONAL
SourceZopNight
Permissions usedcloudsql.instances.list · cloudsql.instances.get

What a zonal Cloud SQL instance loses in a zone outage

A Cloud SQL instance configured for high availability is a regional instance: a primary in one zone and a standby in another, kept in step by synchronous replication to each zone’s persistent disk. Google’s HA overview describes how, if the primary or its zone fails, the standby becomes the new primary and clients follow it through the same static IP address. That switch is a failover.

A zonal instance has none of this. When its zone has trouble, the database is simply unavailable, and recovery means waiting for the zone or restoring a backup into a new instance somewhere else. For a test database that is fine. For the database behind a checkout, a login service or an API, it is a single point of failure that no amount of application-level retry can hide.

Checking availability type across instances

Terminal window
gcloud sql instances list \
--format="table(name, region, settings.availabilityType, settings.userLabels)"

REGIONAL means HA is on; ZONAL means it is not. The labels column helps you check which instances your own tagging marks as production.

How ZopNight decides an instance is production and unprotected

Two conditions must both hold.

  1. The instance is treated as production. Either its name contains prod, production or live (case does not matter), or it carries an env, environment, stage or tier label whose value is prod, production, prd or live. A name containing only prd does not count.
  2. The instance lacks HA. ZopNight reads the HA posture Cloud SQL reports for the instance and fires when the availability type is present and is anything other than REGIONAL.

There is no metric or time window. The check reruns on every evaluation.

Instances this rule does not report

Non-production instances are never flagged, however they are configured: a zonal dev database is a reasonable choice, not a finding. If ZopNight cannot tell what availability type an instance has, it does not guess that HA is missing. A regional instance is always silent, even if you have not added a replica or tested failover.

A reliability gap, not a saving

This rule carries no saving; fixing it adds cost. Google states plainly that an HA-configured instance costs twice as much as a standalone one, covering CPU, RAM and storage, so decide per database whether the outage risk justifies doubling its bill. For a revenue-bearing production database it almost always does.

Turning on high availability

  1. Confirm the instance has automated backups and binary logging where required, since the HA configuration steps pass both.

  2. Convert it:

    Terminal window
    gcloud sql instances patch INSTANCE_NAME \
    --availability-type REGIONAL \
    --enable-bin-log \
    --backup-start-time=HH:MM
  3. Plan for a restart. Google notes that the change restarts the instance, usually within a few minutes, but up to an hour for a large disk or heavy load, and that it cannot be stopped once started.

  4. After it finishes, test how your application reconnects by triggering a failover in a quiet period.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·