# The Event Readiness Playbook: Explained

> A practical guide to pre-scaling cloud infrastructure for sales, launches, and traffic spikes using cloud-native scheduled actions across AWS ASG, ECS, GCP MIG, and Azure VMSS.

Source: https://zop.dev/learn/event-readiness-playbook
Published: 2026-07-01 · Author: avinash-gaurav · Tags: zopnight, learn

---

Black Friday breaks more clouds than any other day of the year. Not because the cloud cannot scale, but because teams rely on reactive autoscaling that takes minutes to absorb a 10x burst. The fix is not faster reactive scaling. It is pre-scaling: getting capacity in place before the spike arrives.

Pre-scaling for events has historically been a cron job or a runbook. Engineers wrote shell scripts that bumped ASG capacity at midnight and forgot to roll back at 6 AM the next day. The result: a one-day sale paid for as a one-week scale-out. ZopNight Event Readiness replaces the runbook with a typed lifecycle, cloud-native scheduled actions, and automatic rollback.

This playbook covers the complete pre-scaling workflow: capacity sizing, target selection, the lifecycle from draft to completed, and the rollback contract. It works across AWS ASG, AWS ECS Application Auto Scaling, GCP MIG, and Azure VMSS. Databases attach as monitor-only because DB tier changes belong with your DBA team.

This guide keeps the theory short and spends most of its length on what you can actually do. Every recommendation here is one ZopNight can help you execute, starting from a read-only connection.

## Sizing the event

The capacity engine accepts a multiplier (e.g., 5x baseline) or expectedRequests. Multiplier is fast and works when you have prior event data. expectedRequests is precise when you have load test results. Either input produces per-target scaledMin, scaledMax, and scaledDesired. The wizard step 2 preview lets you override calculator output before persisting. The cost estimate carries an isEstimated badge and pulls from the aggregator pricing cache. The DB impact preview computes connection-pool math from current resource configuration, no hardcoded specs.

## Picking targets

ASG, ECS service, MIG, and VMSS are the supported scaling targets. Each must already have an autoscaler policy in ZopNight or a pre-existing cloud policy. Databases attach as monitor-only, ZopNight surfaces sizing recommendations from connection-pool math but never modifies customer DBs. DB tier changes and replica scaling stay with your DBA team. The Event Readiness wizard exposes a calculate-db endpoint that computes a recommendation from the aggregator pricing cache.

## The lifecycle and rollback

States: draft, scheduling, scheduled, scaling_up, active, scaling_down, completed. Cancelled is terminal. Concurrent schedule requests on the same event are rejected so the queue never fights itself. The executor writes cloud-native scheduled actions (no polling): AWS ASG PutScheduledUpdateGroupAction, AWS ECS PutScheduledAction, GCP scalingSchedules, Azure FixedDate profile. originalMin/Max/Desired snapshotted on schedule. Rollback restores them on complete or cancel. Cancel sweeps every zopnight-event-* scheduled action so partial-failure does not leak orphans.

## Provider quirks

GCP cancel uses Autoscalers.Update with ForceSendFields:["ScalingSchedules"] because PATCH merges map keys and silently leaves stale entries. AWS ECS joins per-action errors via errors.Join so partial-failure messages name every action that failed to delete. Azure VMSS rejects latency, error_rate, and queue_depth triggers without an explicit metric_config because cross-resource metrics need a target URI. These are not edge cases, they are the difference between a clean cancel and a stuck schedule.

## Key takeaways

- Pre-scaling beats reactive autoscaling for known events because cloud-native scheduled actions commit capacity before the spike.
- ZopNight Event Readiness covers ASG, ECS, MIG, and VMSS with one wizard and one rollback contract.
- Databases attach as monitor-only, ZopNight surfaces sizing recommendations but never modifies customer DBs.
- Deterministic scheduleIds keyed on eventId mean re-queue is a no-op and cancel cleanly removes every schedule.

## Where ZopNight fits

ZopNight turns this from reading into doing. It ships 490 built-in audit rules across AWS (216), GCP (127), and Azure (147), 124 of those recommendations are wired to act end to end, 28 one-click and 96 guided, and it starts read-only so you can see the opportunity before you act on any of it. The most direct place to begin is scheduling non-production resources to your working hours, which is covered in the [FinOps](https://zop.dev/learn/finops) guide and shown concretely for [AWS EC2](https://zop.dev/zopnight/aws/ec2).

## How ZopNight schedules non-production resources

The loop that does this is deliberately mechanical, and it starts read-only. You connect your cloud provider with a read-only role, and ZopNight discovers every non-production resources across your regions and accounts. It records a per-action permission verdict for each one, so you can see where it can list a resource but not yet stop it, and you review that inventory, filter it by status or type, and search for the specific resources you care about before anything is scheduled.

Scheduling itself is a cron you write once in plain terms, stop at 7 PM, start at 8 AM on weekdays, pinned to your timezone so the jobs fire at local business hours rather than UTC. A weekly 24-hour grid shows the schedule visually so you catch gaps and overlaps before you save, and an estimate of active versus inactive hours appears before you commit. Resources attach individually or bundle into groups like "dev-cluster" or "staging-db" so a whole environment follows one cadence.

Actions run in dependency order, so a database comes up before the app server that depends on it. When something needs to stay up, an override forces a non-production resources ON or OFF for a defined window, carries a reason so teammates understand why it exists, and expires automatically so nothing is left running by accident. If a start or stop fails, ZopNight retries up to three times and falls back to a dead-letter queue rather than silently dropping the action, and every state change lands in an audit trail that records whether a schedule, an override, or a specific user triggered it.

## Getting started

Getting started is intentionally low-stakes:

- Connect your cloud provider with a read-only role. Nothing is scheduled or changed at this stage.
- Let ZopNight discover your non-production resources and review exactly what it found, filtered by account, region, and status.
- Create a schedule in your timezone and attach the non-production resources or groups you want it to cover.
- Watch the first cycle run, with Slack, Teams, or Google Chat notifications on every start, stop, and failure, then layer in idle cleanup and guided rightsizing.

Production stays excluded by default throughout, and because discovery and recommendations are read-only, you can prove the value before you enable a single action.

## Frequently asked questions

### How do I size for a launch with no prior data?

Start with a multiplier from a similar event in the past, even an internal load test. The wizard step 2 preview lets you override before persisting. Plan for 1.5x your best estimate to leave headroom.

### Can I cancel mid-event?

Yes. Cancel is idempotent. Deterministic scheduleIds mean cancel cleanly removes the schedules from every target. Notifications fire on cancel and the audit log captures the cancel reason.
