Skip to main content
zopnightgcpkubernetes

Schedule GKE StatefulSets: Stop Stateful Workloads Safely at Night

Most of what teams spend on GKE StatefulSets in non-production is spent while nobody is watching. In a development, staging, or QA account these resources are billed for every hour they exist, but the people who use them work a fraction of those hours. Nights, weekends, holidays, and the long tail of “we’ll get back to that environment next sprint” all meter at full price.

ZopNight closes that gap by scheduling your GKE StatefulSets to run on your team’s hours and stop the rest of the time, without touching your data and without a migration. It is the most direct lever in FinOps, and for teams comparing tools it is where ZopNight pulls ahead of dashboard-first platforms like CloudHealth.

Why GKE StatefulSets cost more than they should

Cloud providers bill GKE StatefulSets by the hour whether or not anyone is using them, and most non-production fleets default to running 24/7 out of habit rather than need. At typical on-demand rates ($0.10–$5.00/hr) an always-on instance quietly bills the same on a Tuesday afternoon as it does at 3 AM on a Sunday.

Here is the shape of it. Say a box runs at a typical on-demand rate of $0.17 an hour. Left on around the clock that is about $124 a month, because a month is roughly 730 hours. Confine it to a single-shift work week, about 50 hours, and you pay for 50 hours instead of 730. Your own rate and hours will differ, but the ratio is the point: in non-production, most of the meter runs while nobody is working.

The instinct is usually to reach for a smaller instance type. But rightsizing only helps a resource that is genuinely too big; it does nothing for a correctly-sized resource that simply runs when no one is around. The larger, easier win is refusing to pay for the hours nobody is working, and it carries none of the performance risk of down-sizing a box that might spike tomorrow.

Stopping and starting GKE StatefulSets safely

ZopNight handles the Scale replicas to zero via Kubernetes API (ordered termination) on every GKE StatefulSets in scope, in dependency order so nothing comes up before what it depends on. A scheduled stop preserves your data exactly as a normal power-off would; ZopNight never terminates or deletes the resource, and idle detection watches CPU, network, and disk signals to surface the GKE StatefulSets that are running but doing nothing.

In practice the setup is: Connect your GCP project and GKE cluster. ZopNight discovers all StatefulSets; Select StatefulSets to schedule by namespace, label, or name pattern; ZopNight scales replicas to zero in reverse ordinal order. PVCs preserved; Before business hours, ZopNight restores replicas in ordinal order, waiting for readiness.

If you run GKE StatefulSets you probably also run GKE, GKE Deployments, Cloud SQL, and scheduling them together is where the dependency ordering earns its keep. Related reading: scheduling and cost optimization.

How ZopNight schedules GKE StatefulSets

The loop that does this is deliberately mechanical, and it starts read-only. You connect GCP with a read-only role, and ZopNight discovers every GKE StatefulSets across your regions and accounts. It records a per-action permission verdict for each one, so you can see where it can list a resource but not yet stop it, and you review that inventory, filter it by status or type, and search for the specific resources you care about before anything is scheduled.

Scheduling itself is a cron you write once in plain terms, stop at 7 PM, start at 8 AM on weekdays, pinned to your timezone so the jobs fire at local business hours rather than UTC. A weekly 24-hour grid shows the schedule visually so you catch gaps and overlaps before you save, and an estimate of active versus inactive hours appears before you commit. Resources attach individually or bundle into groups like “dev-cluster” or “staging-db” so a whole environment follows one cadence.

Actions run in dependency order, so a database comes up before the app server that depends on it. When something needs to stay up, an override forces a GKE StatefulSets ON or OFF for a defined window, carries a reason so teammates understand why it exists, and expires automatically so nothing is left running by accident. If a start or stop fails, ZopNight retries up to three times and falls back to a dead-letter queue rather than silently dropping the action, and every state change lands in an audit trail that records whether a schedule, an override, or a specific user triggered it.

Getting started

Getting started is intentionally low-stakes:

  • Connect GCP with a read-only role. Nothing is scheduled or changed at this stage.
  • Let ZopNight discover your GKE StatefulSets and review exactly what it found, filtered by account, region, and status.
  • Create a schedule in your timezone and attach the non-production resources or groups you want it to cover.
  • Watch the first cycle run, with Slack, Teams, or Google Chat notifications on every start, stop, and failure, then layer in idle cleanup and guided rightsizing.

Production stays excluded by default throughout, and because discovery and recommendations are read-only, you can prove the value before you enable a single action.

faq

Questions we get a lot.

If yours isn't here, email us and we'll answer directly.

Does scaling a GKE StatefulSet to zero lose my data?

No. PersistentVolumeClaims and their backing GCE persistent disks are preserved when pods are terminated. Data remains intact and pods reattach to the same PVCs on scale-up.

How does ZopNight handle ordered termination on GKE?

ZopNight respects StatefulSet ordering. Pods terminate in reverse ordinal order and start in forward order, maintaining data consistency for databases and distributed systems.

Can I take disk snapshots before scaling down?

Yes. ZopNight can create GCE persistent disk snapshots of PVCs before scaling down StatefulSets. This provides a point-in-time recovery option in addition to the preserved PVCs.

What about StatefulSets in GKE Autopilot?

ZopNight works with GKE Autopilot StatefulSets. Scaling to zero releases the underlying Autopilot-managed node resources, and GKE allocates new resources when pods scale back up.

Does this work with Cloud SQL proxy sidecars?

Yes. Pods with Cloud SQL proxy sidecars are handled normally. The proxy sidecar terminates with the pod and reconnects when the pod restarts. Ensure your Cloud SQL instance is running before the StatefulSet starts.

Stop watching the waste.
Start cutting it.

See. Find. Fix. Automatic.

Connect your first cloud account in under 5 minutes. See your first remediation in under 7. No credit card required.

CDCR connect detect classify remediate
full audit every action traceable
read-only default access
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·