# Cloud Cost Playbook: SRE Teams

> How SRE Teams cut non-production cloud waste across AWS, GCP, and Azure with ZopNight. Read-only setup, scheduling first, production excluded by default.

Source: https://zop.dev/for/sre-teams
Published: 2026-07-01 · Author: avinash-gaurav · Tags: zopnight, role, sre-teams

---

ZopNight smart tags production environments and protects them from scheduling. Non-production environments get stopped and started on schedule with dependency ordering and anomaly detection.

If you are a SRE Team, the cloud bill is partly yours to answer for, and non-production waste is the part you can move fastest. The work below is ordered by leverage: the safest, highest-impact action first, then progressively more targeted ones.

## What tends to hurt SRE Teams

The recurring pressure points:

- Risk of accidentally stopping production resources.
- No visibility into non-production resource costs.
- Manual start/stop processes are error-prone.
- Anomaly detection for unexpected resource state changes.

None of these is a process failure; they are what happens when a team ships quickly and non-production keeps running around the clock, the same as production.

## The playbook

Start by scheduling non-production across AWS, GCP, and Azure to your working hours, the fastest and most reversible win. Then add idle detection for the resources that are unused even during the day, and guided rightsizing for the ones that are genuinely oversized. Finally, attribute what remains with showback so every team sees its own spend.

ZopNight runs this loop for you:

- Production safeguards, smart tagged and protected.
- Anomaly detection alerts on unexpected state changes.
- Dependency-aware sequencing prevents boot order issues.
- Full audit trail for compliance and incident review.

It ships 490 built-in audit rules across AWS (216), GCP (127), and Azure (147) and starts read-only, so you can stand it up without a change-management fight. See the same approach applied to [AWS EC2](https://zop.dev/zopnight/aws/ec2) and the discipline behind it in [FinOps](https://zop.dev/learn/finops).

## SRE Teams cloud cost by provider

The same role playbook, one cloud at a time: [AWS](https://zop.dev/for/sre-teams/aws), [GCP](https://zop.dev/for/sre-teams/gcp), [Azure](https://zop.dev/for/sre-teams/azure).

## How ZopNight schedules non-production resources

The loop that does this is deliberately mechanical, and it starts read-only. You connect AWS, GCP, and Azure with a read-only role, and ZopNight discovers every non-production resources across your regions and accounts. It records a per-action permission verdict for each one, so you can see where it can list a resource but not yet stop it, and you review that inventory, filter it by status or type, and search for the specific resources you care about before anything is scheduled.

Scheduling itself is a cron you write once in plain terms, stop at 7 PM, start at 8 AM on weekdays, pinned to your timezone so the jobs fire at local business hours rather than UTC. A weekly 24-hour grid shows the schedule visually so you catch gaps and overlaps before you save, and an estimate of active versus inactive hours appears before you commit. Resources attach individually or bundle into groups like "dev-cluster" or "staging-db" so a whole environment follows one cadence.

Actions run in dependency order, so a database comes up before the app server that depends on it. When something needs to stay up, an override forces a non-production resources ON or OFF for a defined window, carries a reason so teammates understand why it exists, and expires automatically so nothing is left running by accident. If a start or stop fails, ZopNight retries up to three times and falls back to a dead-letter queue rather than silently dropping the action, and every state change lands in an audit trail that records whether a schedule, an override, or a specific user triggered it.

## Getting started

Getting started is intentionally low-stakes:

- Connect AWS, GCP, and Azure with a read-only role. Nothing is scheduled or changed at this stage.
- Let ZopNight discover your non-production resources and review exactly what it found, filtered by account, region, and status.
- Create a schedule in your timezone and attach the non-production resources or groups you want it to cover.
- Watch the first cycle run, with Slack, Teams, or Google Chat notifications on every start, stop, and failure, then layer in idle cleanup and guided rightsizing.

Production stays excluded by default throughout, and because discovery and recommendations are read-only, you can prove the value before you enable a single action.

## Frequently asked questions

### Where should a SRE Team start?

With scheduling non-production across AWS, GCP, and Azure to business hours. It is reversible, needs only a read-only role, and usually shows results in the first cycle.

### Will this touch production?

No. Production is excluded by default; only the non-production resources you choose are scheduled.
