Skip to main content Skip to content

Recommendation rules

Browse the 650+ audit rules by category to see what ZopNight catches on AWS, GCP, Azure, Kubernetes, Databricks, Snowflake, and Microsoft 365.

13 min read

ZopNight ships with 650+ audit rules in 11 categories, covering every major resource type on AWS, GCP, and Azure, plus Kubernetes, Databricks, Snowflake, and Microsoft 365. Use this page to see what each category catches before you connect an account, or to find the rule behind a finding. The full per-rule detail (evidence shape, math, remediation steps) is visible inside each recommendation card in the app.

Recommendations Savings page switched to Rule view, each row naming the rule, its category and ID, an Auto-remediation badge where one applies, the monthly saving, and how many resources it fired on, with the filter panel open on cloud account selection

Recommendations → Savings in Rule view: every rule that fired, what it is worth, and how many resources it caught. Filters narrow by account, type, cost, date, viewed state, and lifecycle.

Coverage at a glance

650+ rules in total. Alongside the cloud-native rules for each provider, the catalogue includes these families:

FamilyRulesNotes
Kubernetes workloads43 per provider, plus a posture familySame rules on EKS, GKE, and AKS; the posture family adds HPA, PodDisruptionBudget, quota, node-pressure, networking, and storage checks
Databricks15 per cloudSame rules on Azure, AWS, and GCP Databricks workspaces
Snowflake20Warehouses, tables, materialized views, pipes, stages, and the account. See Snowflake
Autoscaler5 activeScaling configuration on ASG and VMSS
Microsoft 3651Unused Microsoft 365 Copilot seats. See Microsoft 365

All of these families are counted inside the 650+ total, not added on top of it.

Idle

Resources running but doing no useful work. This is the largest savings category for most accounts.

Sample rules:

Resource typeWhat it catches
EC2Instance stopped for 30 days or more that still bills for its attached EBS volumes and Elastic IPs
LambdaFunction averaging under 1 invocation a day over 30 days
NAT GatewayNAT gateway with no connections and no bytes in any direction over 30 days
Load BalancerLoad balancer with no traffic over 30 days and no registered targets
ECS serviceRunning service with no traffic and near-idle CPU and memory over 30 days
EFSFile system with no reads or writes over 30 days
OpenSearchDomain with no search or indexing activity and low CPU over 30 days
MemoryDBCluster with no new connections and no read or write commands over 30 days
ElastiCacheCluster with no connections, no cache hits or misses, and low CPU over 30 days
RDSInstance with no database connections
Transit GatewayTransit gateway with no bytes in or out over 30 days
VPC Interface EndpointInterface endpoint processing under 1 MiB over 30 days
EKSAbandoned node group whose leftover volumes and load balancers still bill
Microsoft 365 CopilotAssigned Copilot seat its holder has not used for 3 or more consecutive months

An ECS service scaled to zero (desired and running both 0) is raised separately as an advisory cleanup with no dollar figure, because it no longer carries compute cost of its own.

Plus equivalents on GCP and Azure: idle Azure VMs, Azure SQL databases, and storage accounts; idle GCP Spanner, Bigtable, and Memorystore instances; idle GCP global load balancers; and so on.

Rightsizing

Resources sized larger than the workload needs. Dropping one or two instance sizes lowers the bill and keeps the same SLO.

Sample rules:

Resource typeWhat it catches
EC230-day average CPU under 5% (one size down) or under 2% (two sizes down); names the specific smaller instance type
EBSgp2 to gp3 migration (cheaper, often faster)
RDSInstance with CPU under 10% over 30 days; one size down
LambdaPeak memory used under 50% of the configured memory
EBSProvisioned IOPS over-allocation
ECS FargateOversized task definition
GlueJob DPU rightsizing
VPC Flow LogsFlow logs written to CloudWatch Logs that would cost less in S3
GKENode-pool machine size, when average node CPU is under 50%
AKSAgent-pool VM size, when peak node CPU is under 40% and memory under 50%

K8s workload rightsizing rules also catch over-provisioned CPU and memory requests using the HPA ScalingLimited signal.

Schedule

Resources running 24×7 that have a clear off-hours usage pattern. These are strong candidates for ZopNight’s start/stop scheduling.

Sample rules:

Resource typeWhat it catches
RedshiftCluster with no database connections over 30 days; recommends pausing it
GCP VMIdle or non-production VM with a measured off-hours window
Cloud SQLInstance with a measured idle window in its connections
Azure VMNon-production VM (by environment tag, name, or resource group) with an off-hours idle pattern
Azure VMDev/test VM without auto-shutdown
Azure VM Scale SetScale set with a measured idle window
Azure SQLDev/test database that could use the serverless tier and auto-pause
Vertex AI WorkbenchNotebook with no idle shutdown configured

Schedule recommendations integrate with ZopNight’s scheduling feature; Review & remediate opens schedule creation with the resource prefilled.

Discount

Reserved Instance, Savings Plan, and Committed Use opportunities: places where steady, predictable usage justifies a commitment.

Sample rules:

Resource typeWhat it catches
EC2Reserved Instance opportunity: at least 60 days of history, running through the month, and at least 30% CPU or memory utilisation
RDSReserved Instance opportunity
EC2Compute Savings Plan opportunity for production workloads
EC2Spot adoption candidate
ElastiCacheReserved node opportunity for production caches

Plus GCP Committed Use Discount and Azure Reservation equivalents. Commitment savings are priced from the provider’s real 1-year and 3-year No-Upfront rates; see Commitment recommendations.

Orphan

Resources no longer attached to anything that needs them. These are often the easiest wins, because there’s nothing to break.

Sample rules:

Resource typeWhat it catches
EBSUnattached EBS volume
Elastic IPUnassociated Elastic IP
EBSOrphaned EBS snapshot (no source volume)
ECRImages not pulled in over 90 days
ECRRepository never pulled over 90 days
GCPUnattached persistent disk
VPCUnused Internet Gateway
Azure Network WatcherDuplicate flow logging, or a flow log on a gateway subnet
Azure Network WatcherConnection monitor with no surviving source

K8s orphan rules also catch unbound PVCs, released PVs, and Services with no endpoints.

Compliance

Configuration drift from best practice. Not every rule here is directly about cost. Most cover operational hygiene that compounds into cost and risk over time.

Sample rules:

Resource typeWhat it catches
RDSInstance not using Multi-AZ
RDSInstance publicly accessible
IAMRole with wildcard principal
IAMRole with AdministratorAccess
IAMAccess key older than 90 days
IAMUser without MFA enabled
Security GroupUnrestricted inbound access
S3Bucket with public access
VariousResource not encrypted at rest
EC2IMDSv2 not required
EC2EBS optimisation not enabled (older instance families)
EC2Detailed monitoring not enabled
LambdaUsing a deprecated runtime

Security

Risky configurations with operational implications. ZopNight lists them alongside cost recommendations so the same audit pass catches both.

Sample rules:

Resource typeWhat it catches
KubernetesPrivileged container
KubernetesContainer running as root
KubernetesContainer with host network
KubernetesIngress without TLS
KubernetesService exposed as NodePort or LoadBalancer
Azure StorageAccount not using customer-managed key encryption
Vertex AI WorkbenchNotebook with Secure Boot disabled

Reliability

Production-shape issues that cause incidents and rework. Mostly Kubernetes workload rules.

Sample patterns:

  • Single-replica deployment
  • HPA at max replicas (scaling-limited)
  • Missing readiness/liveness probes
  • Missing requests/limits
  • Degraded workloads (deployments, statefulsets, or daemonsets not fully ready)
  • Failed jobs

Kubernetes idle rules separately catch stopped deployments, suspended cronjobs, and cronjobs that never ran.

Performance, governance, and advisory

Three smaller categories sit alongside the eight above:

CategoryWhat it covers
PerformanceA measured bottleneck worth clearing: undersized storage tiers, throttled operations, a workload sustained near its limit. Where a concrete target can be named, the finding asks you to resize; otherwise it asks you to tune.
GovernanceMissing tags and labels that break cost attribution: no environment, team, or cost-centre value where your policy expects one.
AdvisoryFindings worth knowing about that carry no dollar figure and no one-click action. They render as a numbered playbook rather than a savings number.

These three never carry a computed severity; they keep the severity their rule declares. See Recommendations.

Autoscaler

Rules specific to scaling configuration. Six are defined and five are active today.

What it catchesProviders
Scaling group has no scaling policiesAWS (ASG)
Autoscale setting missingAzure (VMSS)
Scaling target too high (> 90%)Azure (VMSS)
Scaling target too low (< 30%)Azure (VMSS)
Cooldown too short (< 120s)Azure (VMSS)

The GCP MIG rule (autoscaler missing) is defined but not active: a MIG’s spend is already counted on its member VMs, so a group-level finding would double-count it.

These rules are actionable via one-click apply through the autoscaler, provided the account is in Recommend or Autopilot mode; Monitor mode is read-only and can’t apply. See Autoscaling.

Databricks

A provider-parameterised rule family fires across all three clouds (15 rules per cloud) for Databricks compute: clusters, SQL warehouses, instance pools, jobs, and model-serving endpoints. See Databricks for the connection model.

The family covers, among others:

What it catchesApplies to
Cluster missing auto-terminationAll-purpose clusters
SQL warehouse without auto-stopSQL warehouses
Model-serving endpoint left always-onServing endpoints
Instance pool holding too many idle VMs (min-idle too high)Instance pools
Cluster oversized for its workloadClusters
Autoscaling disabled on an eligible clusterClusters
Photon disabled where it would cut runtime costClusters
Job running on an expensive all-purpose clusterJobs
On-demand workers that could be spotClusters
No cluster policy attachedClusters
Missing cost-allocation tagsClusters, warehouses, jobs
Orphaned job (no recent runs)Jobs
Orphaned SQL warehouseSQL warehouses

Snowflake

20 rules covering Snowflake warehouses, tables, materialized views, pipes, stages, and the account itself. Snowflake connects as its own account (the cloud it runs on is read from the account host); its rules are part of the 650+ total above. For connecting and scheduling Snowflake, see Snowflake.

What it catchesApplies to
Auto-suspend unset, at zero, or over 120 secondsWarehouses
Idle warehouse: credits burned with zero queries over 14 daysWarehouses
Minimum cluster count set above what the load needsWarehouses
Statement timeout unset or left at the 48-hour defaultWarehouses
Oversized: p95 load under 30% with no remote spillWarehouses
Maximum cluster count above the observed peakWarehouses
Undersized: heavy remote spillageWarehouses
Under-utilised: active under 4 hours a day over 30 daysWarehouses
No resource monitor attachedWarehouses
Missing team / cost-centre / environment tagsWarehouses
Time Travel retention longer than needed on a large tableTables
Permanent table carrying a large Fail-safe costTables
Large cold table: over 100 GiB with no reads in 90 days (needs 90 days of history; not raising findings yet)Tables
Auto-clustering costing more than the queries justifyTables
Search optimization costing more than the lookups justifyTables
Materialized view whose refresh cost exceeds its query benefitMaterialized views
Snowpipe ingesting many small filesPipes
Stale stage with no COPY in over 30 daysStages
Steady baseline spend: consider Capacity pricing (needs 90 days of history; not raising findings yet)Account
Premium edition worth reviewing for downgradeAccount

Five of these carry a dollar figure, priced from your credit consumption at your edition’s rate: the idle and oversized warehouse rules, and the auto-clustering, search-optimization, and materialized-view refresh rules. The rest are configuration posture, where there is no computable saving, so they carry no number rather than a fabricated one.

Severity

Severity is assigned per finding, not per rule. The same rule can fire as Critical on one resource and Medium on another, depending on the size of the saving relative to your spend. For cost findings the score is computed against your own organisation’s spend; see How severity is set.

SeverityRough meaning
CriticalLarge savings relative to your spend, or a severe security or compliance finding
HighSubstantial savings, or a clear best-practice violation
MediumA moderate saving relative to your spend, or a moderate configuration finding
LowSmall absolute savings, or a low-confidence signal

To prioritise, start with Critical findings and work down.

Next steps

Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·