Skip to main content Skip to content

Cost anomalies

Find and investigate daily cost spikes across seven dimensions, from the whole organisation down to a single resource, and route them to the channels you choose.

7 min read

A “surprise bill” is rarely a single bad cloud day. It’s usually a small spike that compounded for a week before anyone noticed. ZopNight’s anomaly detector runs daily, compares every dimension against its own baseline, and surfaces deviations the day they happen. Every flag comes with a root cause analysis so the investigation starts at “why” instead of “where.”

Cost Anomaly drawer for a critical flag: actual spend against the expected forecast with an 85 percent deviation, three plain-language insights naming the cause, a per-resource-type breakdown of daily cost change, and the three affected resources with type, region, and daily cost

Cost Anomaly drawer; what was expected, what actually landed, why it moved, and which resources carried it.

Before you start

  • Cost history. Each detection method needs at least 4 data points and a baseline of at least $1/day before it flags, so a newly connected account starts flagging after a few days of cost records.
  • Teams and resource groups, if you want anomalies at those levels. See Resource groups.
  • Optional: a notification channel. Detection runs without one; subscribe a channel only if you want to be alerted. See Get notified.

How it works

The anomaly cron runs daily at 20:55 UTC. Detection always evaluates and stores anomalies for every organisation, whether or not you have a notification channel configured; notification is a separate, optional layer on top.

Seven dimensions watched

Anomaly detection runs against each of these every day:

DimensionWhat “anomalous” means at this level
Org-wideTotal spend deviates from the 7-day rolling average
Cloud accountOne account’s spend deviates from its own 7-day average
Resource typeSpend on one resource type deviates from its own average
Resource groupA group’s spend deviates from its own average
ResourceAn individual resource’s daily cost deviates; capped to the top 10 per org per day to avoid noise
TeamA team’s allocated spend deviates; with redistribution suppression (see below)
Azure resource group / tenantAn Azure resource group’s or Azure tenant’s spend deviates from its own average

Detection methods

Two methods run in parallel; the higher severity wins.

MethodWhat it computes
Percent deviationToday’s value vs the 7-day rolling average, as a percentage
Z-scoreToday’s value’s deviation from the mean in standard deviations

Both methods need at least 4 data points and a minimum $1/day baseline before they flag. The stddev must be at least 10% of the mean. This filters out flapping noise on very stable baselines.

Severity bands

SeverityThreshold (percent deviation)
Warning30%–150% above baseline
Critical150%–500% above baseline
EmergencyMore than 500% above baseline

Because the z-score test runs too and the higher severity wins, an anomaly can land in a higher band than its percent deviation alone would give: for example, Critical at +85% on a series that normally barely moves.

Root cause analysis

Every flagged anomaly carries an RCA tag identifying the most likely cause:

Instance resize

The hourly rate on a resource changed. Usually means someone upsized.

New resource

A resource that didn’t exist a day ago is contributing to today’s spend.

Reservation expiry

A Reserved Instance or Savings Plan term ended. The same workload is now on-demand.

Schedule failure

A ZopNight schedule was supposed to stop the resource but didn’t. Check the resource’s state history for the failed action.

Unscheduled usage increase

The resource’s usage went up without a schedule accounting for it.

The RCA tag is rule-based. Treat it as a starting point: the anomaly drawer shows the underlying data so you can verify.

Team redistribution suppression

A common false positive: shared resources getting reassigned between teams. Team A’s cost spikes and team B’s cost drops by roughly the same amount, so the net change is near zero. ZopNight detects this. If the absolute net change across teams is under 20%, both anomalies are suppressed as a redistribution rather than fired as separate alerts.

Find anomalies

  1. Open the Summary tab

    Go to Costs → Reports → Summary. When anomalies are detected, a banner says how many, and markers on the cost trend chart point at the days they landed.

  2. List them

    Click View All to list the anomalies for the selected date range. The list is paginated on the server.

  3. Open one

    Each anomaly opens the drawer shown above: actual against expected spend, the deviation, plain-language insights, the per-resource-type breakdown, and the affected resources.

Every stored anomaly is available in three places:

  • Cost trend chart: a marker on the day it landed
  • API: GET /orgs/{orgID}/reports/anomalies?from=&to= returns them
  • Anomaly drawer: the full detail, as in the screenshot above

Get notified

Notification is layered on top of evaluation. Subscriptions are per dimension and per severity, so you can route each combination to a different channel:

  • Subscribe channel A to cost.anomaly.team.critical
  • Subscribe channel B to cost.anomaly.resource.emergency

Notifications deduplicate by date and entity, so you don’t get re-paged for the same anomaly on the same day. Severity escalation re-notifies. If an anomaly was Warning yesterday and is Critical today, the Critical notification goes out.

Set up channels and subscriptions under Settings → Notifications → Channels & Alerts.

Performance on large organisations

The cron is adaptive. Small orgs (≤5K resources) run in batches of 25, medium orgs (≤20K) in batches of 5, and large orgs (>20K) one at a time. This caps peak memory.

Inside each org, ZopNight builds an in-memory cost-record index keyed by resource UID. RCA lookups against the index are O(1), so there is no per-resource database round trip when the analyser asks “what was this resource’s rate yesterday?”

Pipeline prefetch overlaps IO with computation: while one org’s anomalies are being computed, the next org’s cost records are being fetched. Group members are fetched once per org and shared across the resource-group and team detectors.

Troubleshooting

A spike I can see was not flagged

Check the guards: at least 4 data points, a baseline of at least $1/day, and a standard deviation of at least 10% of the mean. The resource dimension is capped to the top 10 per organisation per day, and opposite team movements whose net change is under 20% are suppressed as a redistribution.

I was not notified about an anomaly

Notification is optional and per dimension and severity. Confirm a channel is subscribed to that dimension and severity (for example cost.anomaly.team.critical). Repeat notifications for the same entity on the same day are deduplicated; an escalation to a higher severity notifies again.

The root-cause tag looks wrong

The RCA tag is rule-based, a starting point rather than a verdict. Open the anomaly drawer to check the underlying data and the affected resources.

Next steps

Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·