Skip to main content Skip to content

Recommendations

Review ZopNight's evidence-backed findings, understand how each is priced and ranked, and fix, dismiss, or export them.

19 min read

Recommendations are the output of ZopNight’s audit engine: concrete, evidence-backed findings that point at specific resources costing more than they should, or configured less safely than they should be. Each one is a single rule firing against a single resource, with the math, the evidence, and the suggested action all on the card, so you can decide and act without leaving the page.

Recommendations → Savings: Open, Potential Savings, Applied, and Auto-Optimised cards, the System / Custom toggle, status tabs, search and filter controls, category chips, a Resource / Rule toggle, and resource rows with the monthly saving and open count

Recommendations → Savings: headline totals above open findings grouped by resource, each with its monthly saving, plus the category chips and the System / Custom and Resource / Rule toggles.

Before you start

  • A connected cloud account that has finished its first discovery pass. Findings appear once that pass completes. See Cloud accounts.
  • Metric read permissions (CloudWatch, Azure Monitor, Cloud Monitoring). Rules that need metrics cannot fire without them. See Cloud permissions.
  • Billing connected, for billed-cost savings. Without it, savings are priced at rack rate, and on AWS the money-bearing EC2 and RDS rules wait for billed cost before they fire.
  • A Read + Write cloud account, to use Review & remediate. Read-only accounts get every finding but cannot be remediated. See Read-only vs read + write.

How it works

Recommendations are evaluated per resource: each time discovery scans your accounts (about every 6 hours), and whenever a resource is discovered, ZopNight runs the applicable audit rules against it and stores the findings. The Recommendations pages read from those stored findings, so nothing is recomputed when you load a page.

The rules are expensive to run. Some read long metric windows (up to 90 days on AWS CloudWatch, 60 on Azure Monitor, and 42 on GCP Cloud Monitoring; see how far back metrics are read), or a billing-history join, or both. Doing that work at request time would make the page unusable. Doing it ahead of time and storing the results makes the page instant.

Where recommendations live

The Recommendations section of the sidebar has four pages:

PageWhat it holds
SavingsFindings with a dollar figure: idle, rightsizing, orphan, schedule, and discount
ComplianceConfiguration findings against best practice
RiskSecurity posture and single points of failure
AdvisoryFindings with no savings claim, worth knowing about

Each page has a System / Custom toggle (built-in rules versus your own custom recommendations) and a Resource / Rule toggle (group findings by resource, or by the rule that fired).

What’s on a recommendation card

Every card has the same shape regardless of which rule fired.

FieldWhat it is
TitleThe action to take, not the symptom: “Schedule trading-engine-prod-20 to stop during off-hours” rather than “EC2 instance idle, average CPU 4.0%”. Where a fact is missing (no nameable resize target, no measured count), the title falls back to the rule’s own prose rather than half-stating it
ResourceThe cloud resource the rule fired on, with cloud account and region
Current costMonthly cost of the resource as it sits today
Optimised costWhat the cost becomes if you apply the suggested change
Potential savingsThe delta. ZopNight uses real billing data when available, rack rate otherwise
EvidenceThe metric values that triggered the rule (CloudWatch averages, idle hours, etc.)
Why this firedThe measured signal and the window it covered, taken from the evidence the finding actually carries
Cost breakdownWhere a resource bills through attached children (volumes, reserved IPs, cache nodes), each child’s monthly cost and whether acting on the parent stops it
SeverityCritical / High / Medium / Low, with a one-line reason; see How severity is set
ActionSuggested fix. Some actions are one-click via Auto-remediation; others are advisory

How severity is set

For cost findings (idle, rightsizing, orphan, schedule, discount), severity is computed rather than fixed per rule, so the same rule can land as Critical on one resource and Medium on another. The score runs 0 to 100 and buckets as Low (under 40), Medium (40 to 69), High (70 to 89), and Critical (90 or above).

Three inputs drive it:

  • Savings, measured against your own organisation’s daily spend, so a $400/month finding reads differently in a $2k estate than in a $200k one.
  • Confidence, where the rule crossed a metric threshold.
  • Waste ratio, the share of that resource’s own bill the change recovers (applies to rightsizing, discount, and schedule, where it actually varies).

Savings is the signal that decides the band; the other two order findings inside it. Findings that are not cost-driven (security, compliance, governance, reliability, performance, advisory) keep the severity their rule declares, because a flat cost score cannot tell a public bucket from a deprecated runtime.

Every finding shows its severity together with a short reason, so the rating always says why it is worth acting on.

Categories

ZopNight ships 650+ rules in 11 categories. The Kubernetes, Databricks, Snowflake, Autoscaler, and Microsoft 365 rule families are counted inside that total, not added on top of it. See Recommendation rules for the full catalogue grouped by category.

Idle

Resources running but doing no useful work: zero invocations, no connections, no traffic. The biggest single source of waste in most accounts.

Rightsizing

Resources sized for headroom they never use. A smaller instance type runs the same workload for less.

Schedule

Resources running 24×7 that have a clear off-hours usage pattern. These are candidates for ZopNight’s start/stop scheduling.

Discount

Reserved Instance, Savings Plan, and Committed Use opportunities: places where steady usage justifies a discount commitment.

Orphan

Resources no longer attached to anything that needs them: unattached EBS volumes, unassigned Elastic IPs, released PVs.

Compliance

Configuration drift from best practice: public access, missing MFA, unencrypted storage, deprecated runtimes.

Security

Risky workload configurations: privileged containers, containers running as root, host networking, ingress without TLS.

Reliability

Production-shape issues that cause incidents and rework: single-replica deployments, missing requests/limits, missing probes.

Performance

A measured bottleneck worth clearing: throttled operations, undersized tiers, a workload sustained near its limit.

Governance

Missing tags and labels that break cost attribution, such as no environment, team, or cost-centre value.

Advisory

Findings worth knowing about that carry no dollar figure and no one-click action; they render as a numbered playbook.

Review recommendations

Open Recommendations → Savings (or Compliance, Risk, or Advisory), narrow the list, then open a finding to read its evidence.

Filters

Each Recommendations page has the same controls:

  • Tabs: Open, Applied, Dismissed, Auto-Optimised, Bookmarked
  • Category chips: on Savings, All, Discount, Idle, Orphan, Rightsizing, Schedule
  • Filter panel: Account, Type, Cost, Date, Viewed, Lifecycle
  • Search: matches title, resource name, and resource ID
  • Order: default order or a chosen sort

Filters cascade: selecting a cloud account narrows the other options to what that account actually has, so the dropdowns never offer dead ends.

Act on a recommendation

There are three things you can do with any open recommendation:

  1. Apply the suggested change (Remediate)

    For rules on the auto-remediation list, ZopNight can run the change for you. Review & remediate opens a wizard that previews the change, asks for approval if the action is destructive, runs the cloud API call, and validates the result. Some run one-click (Auto); others add a type-to-confirm review (Guided). See Auto-remediation.

  2. Apply it yourself, then mark it resolved

    For advisory rules (cases where ZopNight should not run cloud-side mutations on your behalf), the card shows a numbered “How to fix” playbook. After you’ve done the work in your cloud console or via Terraform, click Mark Resolved so it moves to Applied and drops off the Open list.

  3. Dismiss it

    Sometimes a recommendation doesn’t apply: the resource has a reason to be over-provisioned, or it’s about to be deleted anyway. Dismiss it with a reason and a cooldown (for example Dismiss for 1 month). When the cooldown expires, the recommendation reopens if the resource still matches; if ZopNight verified in the meantime that the change was made, it closes as optimised instead.

Applied, and then verified

Marking a recommendation Applied records what you did. It is not the same as ZopNight confirming it happened, and the two are kept apart deliberately. Status is your claim; verification is what we measured.

After you mark one applied, ZopNight re-reads the resource and reports back:

What we findWhat it means
PendingToo early to tell. We are still inside the billing lag or the window the change has to hold for; a verdict will follow.
AdoptedThe estate changed to match the recommendation.
UnchangedNothing moved. Worth a second look; the change may not have landed.
DivergedSomething changed, but not into the state we asked for.
GoneThe resource no longer exists.
Not verifiableWe have no honest way to check this one, and say so rather than guessing.

What counts as proof depends on the action, not the category. A stop is proven by power state; a schedule is proven by the running fraction dropping over at least a week (not by a status snapshot, which flips depending on when it is sampled); a resize is proven by the instance shape moving; a delete is proven by the resource leaving the inventory. Where a recommendation asks for something none of our data can see (write a policy, review a configuration), it reports Not verifiable instead of inventing a verdict.

A recommendation you dismissed and then fixed anyway is picked up too, and closes as optimised rather than reopening.

Schedule recommendations

A schedule recommendation proposes a recurring stop/start window for a resource that sits idle on a predictable pattern. Its drawer shows a weekly utilisation grid (running versus stopped, hour by hour), the metric chart behind it, and a Today versus After apply strip so you can see exactly which hours the schedule removes.

Schedule recommendation drawer with a weekly schedule grid marking running and stopped hours per day, a CPU utilisation chart, and a Safety and Impact panel comparing today's 24/7 running strip with the after-apply window

Recommendation drawer for a schedule finding; the weekly running pattern, the CPU evidence, and the before and after running window.

ZopNight only raises a schedule recommendation on recent, continuous evidence. It reads the latest 21 days of metrics and requires all of the following:

  • The newest datapoint is no more than 3 days old. This wider allowance replaces the 48-hour freshness check other rules use.
  • The history reaches back at least 18 days.
  • At least 7 distinct days and at least 24 datapoints are observed.
  • No gap in the data is longer than 3 days.

Resources that are simply off part of the time (nights, weekends, a 3-day holiday) still qualify. A resource deliberately kept off for more than 3 consecutive days is not scheduled, because at day granularity that looks the same as a stopped metrics pipeline. If metrics stop arriving, schedule recommendations close on the next pass and reopen when data resumes.

How the saving is priced. Schedule savings are your effective per-running-hour rate multiplied by the running hours the schedule removes. With billing connected, the rate is your last 30 days of billed cost divided by the resource’s measured running hours, so negotiated discounts are included; otherwise it is the hourly catalogue rate. A resource that already runs part-time is priced on the hours it really runs.

When there is no honest saving. If your existing schedule already beats the proposal, or the resource’s running hours cannot be verified, the finding stays visible as an advisory with the derived schedule and the reasoning, and no dollar figure.

Acting on it. Review & remediate on a schedule recommendation opens schedule creation with the resource prefilled, and attaches it when you save. See Scheduling.

Commitment recommendations

Reserved Instance and Committed Use findings are priced from real rates, not a flat discount. ZopNight reads the provider’s published 1-year and 3-year No-Upfront prices and offers one option per term it can price.

A resource must earn a commitment recommendation first: at least 60 days of observed history, running through the month, and at least 30% CPU or memory utilisation. An always-on but under-used box is a rightsizing candidate first.

The math assumes a No-Upfront commitment bills every hour of the term. The saving is measured against what the resource really costs you today, bounded by its measured uptime, and a break-even gate suppresses the option unless your current cost clears the committed price by at least 10%. When a term cannot be priced, the saving shows as unknown rather than a guessed number. Capacity already covered by a reservation or Savings Plan does not re-fire.

Custom recommendations

The built-in catalogue covers the patterns we see across every estate. When you want a threshold that is specific to yours, write it yourself: a watch policy is a recommendation you define from metrics, and its findings appear under the Custom side of the System / Custom toggle on the Recommendations pages.

Custom recommendations list under Governance, Policies, with template chips such as Idle databases and Over-provisioned VMs, and a table of policies showing what each watches, its signals, decision, outcome, and an on/off status toggle

Governance → Policies → Recommendation: watch policies with their scope, signals, decision, outcome, and status, plus starter templates.

You author one in Governance → Policies → Recommendation:

  1. Pick a scope

    A resource selector saying which resources the policy watches.

  2. Define one or more signals

    A signal is a metric, an aggregation, an operator, and a window: avg cpu < 5% over 14 days. Available concepts include CPU, memory, network in/out, disk and IOPS, queue depth, database connections and QPS, invocations, duration, error rate, cache hit rate, uptime, GPU and GPU memory, request units, throttled operations, and message counts. Aggregations are avg, max, min, p95, p99, and sum, where p95 and p99 are percentiles of per-hour peaks and max/min read the per-hour extreme. The window is 1 to 90 days, bounded by how far back each cloud’s metrics are read (90 days on AWS, 60 on Azure, 42 on GCP).

    The wizard greys out metrics your chosen scope cannot collect, so you cannot save a policy that could never fire. A signal with missing data never fires.

  3. Choose the decision and the outcome

    All signals must hold for the policy to fire. Then choose what the finding should recommend (idle, schedule, rightsize, terminate, or a custom label) and the severity you want it to carry.

A custom finding behaves like a built-in one: the same drawer, the same savings basis, the same RBAC, the same $5 floor, the same Dismiss and Mark Resolved actions.

History and bookmarks

Every recommendation carries a History drawer: generated, viewed, bookmarked, applied, dismissed, reopened, and closed, grouped by day. Use it to answer “who dismissed this, and when” without leaving the page. Opening history from a finding narrows to that finding’s rule, so following one thread does not mean scrolling past everything else that happened that day.

Bookmarks are personal. Starring a recommendation keeps it within reach for you and changes nothing for your teammates, and any role that can view recommendations can bookmark one. Your starred findings are on the Bookmarked tab.

Export recommendations

You can export the filtered Recommendations list as an Excel workbook for sharing with finance or the platform team. The export runs as a background job. When the file is ready, a short-lived signed download link appears on the page; for long jobs the link is also emailed, so you can close the tab.

The workbook has two sheets:

  • Executive Summary: who ran it and when, which filters were applied, the headline totals (recommendations, resources affected, potential monthly and annual savings), and breakdowns by recommendation type, cloud provider, service, severity, and status.
  • Recommendation Details: one row per recommendation, with region, resource group, tags, cloud account, monthly and annual savings, the recommended action, the supporting evidence, first detected, age in days, and last evaluated.

Where recommendations show up

Beyond the Recommendations pages, findings also appear here:

  • Resource detail drawer: “Recommendations” tab lists every open recommendation for that resource
  • Costs → Flow: hovering a node shows a “$X reclaimable” callout when the dimension matches a recommendation filter (resource type, cloud account)
  • Dashboard: the Recommendations widget shows top findings by savings
  • MCP server: the AI Assistant can list, filter, dismiss, apply, and schedule recommendations through the MCP server’s recommendation tools (see the MCP setup guide)

Troubleshooting

A recommendation I expected is missing

A few reasons a finding you expect to see isn’t there:

  • Discovery hasn’t reached it yet. A freshly connected cloud account sees its first findings after its first discovery pass completes.
  • Metrics aren’t flowing, or are stale. Rules that need CloudWatch / Cloud Monitoring / Azure Monitor data can’t fire without the right IAM permissions. A rule also stands down when its newest datapoint is more than 48 hours old, rather than reasoning off old evidence; it comes back when fresh data arrives. Schedule recommendations are the one exception: their evidence window allows the newest datapoint to be up to 3 days old (see Schedule recommendations). Check Cloud permissions.
  • The saving can’t be stated concretely. A cost finding must name a real lever and a real dollar figure, or it is not raised. On AWS, money-bearing rules for EC2 and RDS stand down when the resource’s cost is still a rack-rate estimate rather than billed cost, and fire once billed cost arrives.
  • The saving is under $5 a month. Cost findings below $5/month are dropped as noise. Orphan and advisory findings are exempt.
  • A better lever won on the same resource. Where several findings offer mutually exclusive ways to save on one resource (for example Spot versus a Reserved Instance, or two resize targets), only the best one is kept, and the savings on one resource can never add up to more than its cost. An idle finding also suppresses discount findings on the same resource.
  • The schedule evidence isn’t there yet. Schedule recommendations need the 21-day evidence window described in Schedule recommendations.
  • Pricing data is missing. If a rule needs to compute savings but the underlying SKU has no price yet, the rule abstains and the gap is recorded for us to fill. Ask support to backfill.
  • The rule was dismissed previously. Dismissed findings stay dismissed until their cooldown expires. Check the Dismissed tab.
A dismissed recommendation came back

Dismissals carry a cooldown. When it expires, the recommendation reopens if the resource still matches the rule. If ZopNight verified in the meantime that the change was made, it closes as optimised instead. Dismiss again with a longer cooldown if the reason still holds.

An export failed

Exports are persisted, so a crashed export shows up as a Failed row (with reason) rather than disappearing silently. Re-trigger from the UI to retry.

The export's breakdown rows do not add up to the headline

Summary savings are deduplicated per resource: where two findings offer mutually exclusive ways to save on the same resource, only the larger counts. The headline is what you could actually capture, so the breakdown rows can sum to more.

Next steps

Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·