Auto-remediation
Let ZopNight apply a supported recommendation for you, behind a safety gate and admin approval, and confirm the change took effect.
For rules on the auto-remediation list, ZopNight can apply the suggested fix for you, so you don’t switch to the cloud console or copy scripts. Click Review & remediate, walk through the wizard, and ZopNight runs the cloud API call, validates the result, and records every step in the audit log. Rules that aren’t auto-remediable still get a numbered “How to fix” playbook on the same card.
Before you start
- A Read + Write cloud account. Only Read + Write accounts can have resources started, stopped, or remediated; on a Read-only account the change is refused before any cloud call. See Read-only vs read + write.
- The write permission for the action on the cloud side, such as
ec2:StopInstances. See Cloud permissions. - An admin to approve gated operations. When someone other than an admin starts a Pause or Delete, an admin must approve it.
How it works
What’s auto-remediable today
The list covers these families:
| Family | Provider(s) | Flow | What ZopNight does |
|---|---|---|---|
| Start / Stop | AWS, GCP, Azure | Auto | Toggle a resource’s running state. A one-shot stop on a compute instance is always advisory (see below); for off-hours savings, use a schedule |
| Pause | Azure (including Synapse SQL Pool and ML Compute), GCP Dataproc | Guided | Move the resource to its paused billing tier |
| Resize / Rightsize | AWS, GCP, Azure | Mostly Guided | Apply a smaller instance type or scaling configuration |
| Lifecycle setting | AWS, GCP, Azure (incl. Azure Hybrid Benefit) | Mostly Guided | Flip a reversible cost/licensing setting or storage lifecycle policy |
| K8s manifest patch | EKS, GKE, AKS | Guided | Scale a single-replica Deployment 1→2 (never auto; ZopNight never auto-scales a workload you chose as a singleton) |
| Delete (orphan) | AWS, GCP, Azure | Auto or Guided | Remove resources with positive proof that they are unused (see Why some findings stay advisory). Proven orphan deletes run Auto, except snapshot deletes; every other delete is Guided |
A rule leaves the list when it cannot deliver a concrete lever: advisory or $0 findings, rules whose saving is really delivered by an off-hours schedule rather than a one-shot action, and rules folded into a sibling that already fires. The remaining rules render as advisory: the recommendation card shows a “How to fix” playbook with explicit steps for your cloud console or Terraform.
Auto vs Guided
Every rule on the list runs one of two ways:
- Auto: runs end-to-end in one click; the wizard shows progress and you watch. Proven orphan deletes are Auto, except snapshot deletes. (A non-admin initiator still hits an admin-approval gate.)
- Guided: Pause, most Resize, lifecycle, snapshot deletes and every other Delete operation, and every K8s manifest patch. The wizard adds a type-to-confirm modal with an impact preview before executing.
Both flow types render the same button. The difference is what happens after you click it.
Why some findings stay advisory
Whether a rule supports one-click apply and whether it is safe on your particular resource are two different questions, and ZopNight checks both. Every recommendation passes through a safety gate before it earns a Remediate button, so a rule that is auto-remediable in general can still arrive as advisory on a specific resource.
The gate reads only authoritative control-plane fields, never tags and never resource names; those are customer-set and cannot be trusted to decide a destructive action. When the evidence is absent or ambiguous, it fails closed: the finding is surfaced as advisory and a human confirms.
| Action | Gate decision |
|---|---|
| Stop a compute instance | Always advisory. No cloud API tells us what is running inside a generic VM, so a one-shot stop can never be proven safe. Use a recurring schedule instead; a stop is an outage, not a cost lever. |
| Delete a resource | Advisory when the resource is stateful, or when there is no positive proof it is unused. Otherwise allowed. |
| Resize with a restart | Advisory when the resource is stateful. We do not auto-reboot a database. |
| Hot resize, pause, lifecycle setting, Kubernetes manifest patch | Allowed. Reversible or non-destructive. |
Proof that a resource is genuinely unused means one of: an Elastic IP or public IP explicitly not associated; a disk that is unattached and has been detached for at least 7 days; a snapshot whose source volume is provably gone; a backup vault holding zero protected items. Absence of evidence is not evidence, so a resource we simply cannot read stays advisory. A snapshot delete that passes this check still runs Guided, because deleting a snapshot is irreversible and an image may still reference it.
Stateful is decided by what the resource is, not what it is called: any managed data service type (RDS, Aurora, Cloud SQL, ElastiCache, Redshift, DocumentDB, Neptune, MemoryDB, AlloyDB, Memorystore, Bigtable, Spanner, Azure SQL, Cosmos DB, Redis, and the flexible-server variants), anything reporting a database engine, or anything whose tier is reported as data.
Advisory rules
For the rules that aren’t auto-remediable (the balance of the 650+ rule catalogue), the recommendation drawer renders a numbered “How to fix” playbook. Steps are written for the cloud console where possible, with Terraform / CLI variants when they exist.
The playbook is the same content that would have driven an auto-remediation if one existed; the only difference is whether ZopNight runs the call or you do.
Run a remediation
Clicking Review & remediate on any auto-remediable card opens a three- or four-step wizard.
Precondition
ZopNight verifies that the cloud-side state matches what the recommendation assumed. If the resource has changed since the recommendation was computed (manually restarted, terminated, moved to a different account), the wizard surfaces the drift and asks you to refresh the recommendation before proceeding.
Approval (when gated)
When someone other than an admin starts a gated operation such as Pause or Delete, an admin must approve it, in the wizard or from the approval email, before ZopNight runs the cloud API call. An admin who starts the remediation skips this step. Approval links stay valid even if the wizard’s step names change later.
Execute
ZopNight calls the cloud API. The wizard shows the call status in real time. If the call fails, you get a typed error; see Error categories below.
Validate
ZopNight reads the resource back from the cloud API and confirms the change took effect. A successful Validate flips the recommendation to Optimised.
Schedule recommendations are not run through the wizard. Their Review & remediate button opens schedule creation with the resource prefilled, and attaches it when you save. See Scheduling.
Error categories
If the cloud API call fails, ZopNight categorises the error and renders the wizard accordingly.
User action
Yellow. Permission or quota issue. The wizard shows a fix hint (e.g. “Grant ec2:StopInstances on this role”) and a deep link to the cloud console. No retry; fix the cause first.
Transient
Blue. Rate limit (429) or transient 5xx. The wizard offers a Retry button.
System
Red. Unsupported API surface or unknown failure. The wizard shows a copyable diagnostic block and a Contact Support link. No retry from the wizard; needs platform investigation.
Subscription IDs, AWS account IDs, GCP project IDs, and Azure SDK divider lines are stripped from error messages before persistence so audit logs don’t carry sensitive identifiers.
Rollback
For reversible operations, the remediation job’s step-by-step record shows what was changed, so you can restore the previous state. Autoscaler policies go further: ZopNight saves the previous scaling configuration on apply and restores it when you remove the policy. See Autoscaling.
| Operation | Reversible? |
|---|---|
| Start | Yes. Stop |
| Stop | Yes. Start |
| Pause (Azure Synapse SQL Pool, ML Compute) | Yes. Restore |
| Pause (GCP Dataproc) | No. Dataproc pause is terminal; ZopNight surfaces this explicitly before you confirm. |
| Delete | Generally no; undeleting orphan disks is cloud-provider-specific |
Notifications
Remediation sends these notifications:
- Approval-required emails go to the configured admin channel with the subject
[ZopNight] Remediation approval needed. A step is only ever announced once, even if it is retried. - Completed and failed remediations can be sent to any channel subscribed to remediation events in Settings → Notifications → Channels & Alerts. The matching emails are off by default to avoid noise; ask support to enable them if you want them.
Remediation history
Every remediation runs as a job with a step-by-step timeline: kind, subject, status, start/end timestamps, and error category for each step.
The audit log records a row per state transition, including who approved and who rejected. CSV export is available from Activity → Audit Logs in the app.
Troubleshooting
A finding has no Remediate button
Either the rule is not on the auto-remediation list, or the safety gate demoted it on this resource (for example a one-shot stop on a compute instance, or a delete on a stateful resource). The finding keeps its savings and evidence with a “How to fix” playbook instead. See Why some findings stay advisory.
The wizard stops at Precondition
The resource changed since the recommendation was computed (manually restarted, terminated, or moved to a different account). Refresh the recommendation before proceeding.
Execute fails with a yellow User action error
A permission or quota issue on the cloud side. Follow the fix hint (for example, grant ec2:StopInstances on the role), then start again. There is no retry until the cause is fixed.
The remediation is waiting for approval
A non-admin started a gated operation. An admin must approve it, in the wizard or from the [ZopNight] Remediation approval needed email, before the cloud call runs.