Recommendation rules
Browse the 650+ audit rules by category to see what ZopNight catches on AWS, GCP, Azure, Kubernetes, Databricks, Snowflake, and Microsoft 365.
ZopNight ships with 650+ audit rules in 11 categories, covering every major resource type on AWS, GCP, and Azure, plus Kubernetes, Databricks, Snowflake, and Microsoft 365. Use this page to see what each category catches before you connect an account, or to find the rule behind a finding. The full per-rule detail (evidence shape, math, remediation steps) is visible inside each recommendation card in the app.

Recommendations → Savings in Rule view: every rule that fired, what it is worth, and how many resources it caught. Filters narrow by account, type, cost, date, viewed state, and lifecycle.
Coverage at a glance
650+ rules in total. Alongside the cloud-native rules for each provider, the catalogue includes these families:
| Family | Rules | Notes |
|---|---|---|
| Kubernetes workloads | 43 per provider, plus a posture family | Same rules on EKS, GKE, and AKS; the posture family adds HPA, PodDisruptionBudget, quota, node-pressure, networking, and storage checks |
| Databricks | 15 per cloud | Same rules on Azure, AWS, and GCP Databricks workspaces |
| Snowflake | 20 | Warehouses, tables, materialized views, pipes, stages, and the account. See Snowflake |
| Autoscaler | 5 active | Scaling configuration on ASG and VMSS |
| Microsoft 365 | 1 | Unused Microsoft 365 Copilot seats. See Microsoft 365 |
All of these families are counted inside the 650+ total, not added on top of it.
Idle
Resources running but doing no useful work. This is the largest savings category for most accounts.
Sample rules:
| Resource type | What it catches |
|---|---|
| EC2 | Instance stopped for 30 days or more that still bills for its attached EBS volumes and Elastic IPs |
| Lambda | Function averaging under 1 invocation a day over 30 days |
| NAT Gateway | NAT gateway with no connections and no bytes in any direction over 30 days |
| Load Balancer | Load balancer with no traffic over 30 days and no registered targets |
| ECS service | Running service with no traffic and near-idle CPU and memory over 30 days |
| EFS | File system with no reads or writes over 30 days |
| OpenSearch | Domain with no search or indexing activity and low CPU over 30 days |
| MemoryDB | Cluster with no new connections and no read or write commands over 30 days |
| ElastiCache | Cluster with no connections, no cache hits or misses, and low CPU over 30 days |
| RDS | Instance with no database connections |
| Transit Gateway | Transit gateway with no bytes in or out over 30 days |
| VPC Interface Endpoint | Interface endpoint processing under 1 MiB over 30 days |
| EKS | Abandoned node group whose leftover volumes and load balancers still bill |
| Microsoft 365 Copilot | Assigned Copilot seat its holder has not used for 3 or more consecutive months |
An ECS service scaled to zero (desired and running both 0) is raised separately as an advisory cleanup with no dollar figure, because it no longer carries compute cost of its own.
Plus equivalents on GCP and Azure: idle Azure VMs, Azure SQL databases, and storage accounts; idle GCP Spanner, Bigtable, and Memorystore instances; idle GCP global load balancers; and so on.
Rightsizing
Resources sized larger than the workload needs. Dropping one or two instance sizes lowers the bill and keeps the same SLO.
Sample rules:
| Resource type | What it catches |
|---|---|
| EC2 | 30-day average CPU under 5% (one size down) or under 2% (two sizes down); names the specific smaller instance type |
| EBS | gp2 to gp3 migration (cheaper, often faster) |
| RDS | Instance with CPU under 10% over 30 days; one size down |
| Lambda | Peak memory used under 50% of the configured memory |
| EBS | Provisioned IOPS over-allocation |
| ECS Fargate | Oversized task definition |
| Glue | Job DPU rightsizing |
| VPC Flow Logs | Flow logs written to CloudWatch Logs that would cost less in S3 |
| GKE | Node-pool machine size, when average node CPU is under 50% |
| AKS | Agent-pool VM size, when peak node CPU is under 40% and memory under 50% |
K8s workload rightsizing rules also catch over-provisioned CPU and memory requests using the HPA ScalingLimited signal.
Schedule
Resources running 24×7 that have a clear off-hours usage pattern. These are strong candidates for ZopNight’s start/stop scheduling.
Sample rules:
| Resource type | What it catches |
|---|---|
| Redshift | Cluster with no database connections over 30 days; recommends pausing it |
| GCP VM | Idle or non-production VM with a measured off-hours window |
| Cloud SQL | Instance with a measured idle window in its connections |
| Azure VM | Non-production VM (by environment tag, name, or resource group) with an off-hours idle pattern |
| Azure VM | Dev/test VM without auto-shutdown |
| Azure VM Scale Set | Scale set with a measured idle window |
| Azure SQL | Dev/test database that could use the serverless tier and auto-pause |
| Vertex AI Workbench | Notebook with no idle shutdown configured |
Schedule recommendations integrate with ZopNight’s scheduling feature; Review & remediate opens schedule creation with the resource prefilled.
Discount
Reserved Instance, Savings Plan, and Committed Use opportunities: places where steady, predictable usage justifies a commitment.
Sample rules:
| Resource type | What it catches |
|---|---|
| EC2 | Reserved Instance opportunity: at least 60 days of history, running through the month, and at least 30% CPU or memory utilisation |
| RDS | Reserved Instance opportunity |
| EC2 | Compute Savings Plan opportunity for production workloads |
| EC2 | Spot adoption candidate |
| ElastiCache | Reserved node opportunity for production caches |
Plus GCP Committed Use Discount and Azure Reservation equivalents. Commitment savings are priced from the provider’s real 1-year and 3-year No-Upfront rates; see Commitment recommendations.
Orphan
Resources no longer attached to anything that needs them. These are often the easiest wins, because there’s nothing to break.
Sample rules:
| Resource type | What it catches |
|---|---|
| EBS | Unattached EBS volume |
| Elastic IP | Unassociated Elastic IP |
| EBS | Orphaned EBS snapshot (no source volume) |
| ECR | Images not pulled in over 90 days |
| ECR | Repository never pulled over 90 days |
| GCP | Unattached persistent disk |
| VPC | Unused Internet Gateway |
| Azure Network Watcher | Duplicate flow logging, or a flow log on a gateway subnet |
| Azure Network Watcher | Connection monitor with no surviving source |
K8s orphan rules also catch unbound PVCs, released PVs, and Services with no endpoints.
Compliance
Configuration drift from best practice. Not every rule here is directly about cost. Most cover operational hygiene that compounds into cost and risk over time.
Sample rules:
| Resource type | What it catches |
|---|---|
| RDS | Instance not using Multi-AZ |
| RDS | Instance publicly accessible |
| IAM | Role with wildcard principal |
| IAM | Role with AdministratorAccess |
| IAM | Access key older than 90 days |
| IAM | User without MFA enabled |
| Security Group | Unrestricted inbound access |
| S3 | Bucket with public access |
| Various | Resource not encrypted at rest |
| EC2 | IMDSv2 not required |
| EC2 | EBS optimisation not enabled (older instance families) |
| EC2 | Detailed monitoring not enabled |
| Lambda | Using a deprecated runtime |
Security
Risky configurations with operational implications. ZopNight lists them alongside cost recommendations so the same audit pass catches both.
Sample rules:
| Resource type | What it catches |
|---|---|
| Kubernetes | Privileged container |
| Kubernetes | Container running as root |
| Kubernetes | Container with host network |
| Kubernetes | Ingress without TLS |
| Kubernetes | Service exposed as NodePort or LoadBalancer |
| Azure Storage | Account not using customer-managed key encryption |
| Vertex AI Workbench | Notebook with Secure Boot disabled |
Reliability
Production-shape issues that cause incidents and rework. Mostly Kubernetes workload rules.
Sample patterns:
- Single-replica deployment
- HPA at max replicas (scaling-limited)
- Missing readiness/liveness probes
- Missing requests/limits
- Degraded workloads (deployments, statefulsets, or daemonsets not fully ready)
- Failed jobs
Kubernetes idle rules separately catch stopped deployments, suspended cronjobs, and cronjobs that never ran.
Performance, governance, and advisory
Three smaller categories sit alongside the eight above:
| Category | What it covers |
|---|---|
| Performance | A measured bottleneck worth clearing: undersized storage tiers, throttled operations, a workload sustained near its limit. Where a concrete target can be named, the finding asks you to resize; otherwise it asks you to tune. |
| Governance | Missing tags and labels that break cost attribution: no environment, team, or cost-centre value where your policy expects one. |
| Advisory | Findings worth knowing about that carry no dollar figure and no one-click action. They render as a numbered playbook rather than a savings number. |
These three never carry a computed severity; they keep the severity their rule declares. See Recommendations.
Autoscaler
Rules specific to scaling configuration. Six are defined and five are active today.
| What it catches | Providers |
|---|---|
| Scaling group has no scaling policies | AWS (ASG) |
| Autoscale setting missing | Azure (VMSS) |
| Scaling target too high (> 90%) | Azure (VMSS) |
| Scaling target too low (< 30%) | Azure (VMSS) |
| Cooldown too short (< 120s) | Azure (VMSS) |
The GCP MIG rule (autoscaler missing) is defined but not active: a MIG’s spend is already counted on its member VMs, so a group-level finding would double-count it.
These rules are actionable via one-click apply through the autoscaler, provided the account is in Recommend or Autopilot mode; Monitor mode is read-only and can’t apply. See Autoscaling.
Databricks
A provider-parameterised rule family fires across all three clouds (15 rules per cloud) for Databricks compute: clusters, SQL warehouses, instance pools, jobs, and model-serving endpoints. See Databricks for the connection model.
The family covers, among others:
| What it catches | Applies to |
|---|---|
| Cluster missing auto-termination | All-purpose clusters |
| SQL warehouse without auto-stop | SQL warehouses |
| Model-serving endpoint left always-on | Serving endpoints |
| Instance pool holding too many idle VMs (min-idle too high) | Instance pools |
| Cluster oversized for its workload | Clusters |
| Autoscaling disabled on an eligible cluster | Clusters |
| Photon disabled where it would cut runtime cost | Clusters |
| Job running on an expensive all-purpose cluster | Jobs |
| On-demand workers that could be spot | Clusters |
| No cluster policy attached | Clusters |
| Missing cost-allocation tags | Clusters, warehouses, jobs |
| Orphaned job (no recent runs) | Jobs |
| Orphaned SQL warehouse | SQL warehouses |
Snowflake
20 rules covering Snowflake warehouses, tables, materialized views, pipes, stages, and the account itself. Snowflake connects as its own account (the cloud it runs on is read from the account host); its rules are part of the 650+ total above. For connecting and scheduling Snowflake, see Snowflake.
| What it catches | Applies to |
|---|---|
| Auto-suspend unset, at zero, or over 120 seconds | Warehouses |
| Idle warehouse: credits burned with zero queries over 14 days | Warehouses |
| Minimum cluster count set above what the load needs | Warehouses |
| Statement timeout unset or left at the 48-hour default | Warehouses |
| Oversized: p95 load under 30% with no remote spill | Warehouses |
| Maximum cluster count above the observed peak | Warehouses |
| Undersized: heavy remote spillage | Warehouses |
| Under-utilised: active under 4 hours a day over 30 days | Warehouses |
| No resource monitor attached | Warehouses |
| Missing team / cost-centre / environment tags | Warehouses |
| Time Travel retention longer than needed on a large table | Tables |
| Permanent table carrying a large Fail-safe cost | Tables |
| Large cold table: over 100 GiB with no reads in 90 days (needs 90 days of history; not raising findings yet) | Tables |
| Auto-clustering costing more than the queries justify | Tables |
| Search optimization costing more than the lookups justify | Tables |
| Materialized view whose refresh cost exceeds its query benefit | Materialized views |
| Snowpipe ingesting many small files | Pipes |
| Stale stage with no COPY in over 30 days | Stages |
| Steady baseline spend: consider Capacity pricing (needs 90 days of history; not raising findings yet) | Account |
| Premium edition worth reviewing for downgrade | Account |
Five of these carry a dollar figure, priced from your credit consumption at your edition’s rate: the idle and oversized warehouse rules, and the auto-clustering, search-optimization, and materialized-view refresh rules. The rest are configuration posture, where there is no computable saving, so they carry no number rather than a fabricated one.
Severity
Severity is assigned per finding, not per rule. The same rule can fire as Critical on one resource and Medium on another, depending on the size of the saving relative to your spend. For cost findings the score is computed against your own organisation’s spend; see How severity is set.
| Severity | Rough meaning |
|---|---|
| Critical | Large savings relative to your spend, or a severe security or compliance finding |
| High | Substantial savings, or a clear best-practice violation |
| Medium | A moderate saving relative to your spend, or a moderate configuration finding |
| Low | Small absolute savings, or a low-confidence signal |
To prioritise, start with Critical findings and work down.