Azure AI Search services with extra replicas or partitions serving under 0.1 queries per second
What does ZopNight detect here?
Azure AI Search bills an hourly rate per search unit, where units equal replicas multiplied by partitions, whether or not queries arrive. ZopNight flags a service with more than one unit whose `SearchQueriesPerSecond` averaged under 0.1 over 30 days without throttling, and prices the units above your availability floor.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1396 |
| Category | rightsizing |
| Severity | medium |
| Metric | SearchQueriesPerSecond |
| Threshold | average below 0.1 QPS, throttling below 1% |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | Microsoft.Search/searchServices/read · Microsoft.Insights/Metrics/Read |
Where it applies
Search units bill whether anyone searches or not
For provisioned tiers, Microsoft’s capacity planning guide says you pay an hourly rate measured in search units regardless of usage, and that the unit count is replicas multiplied by partitions. Two replicas and two partitions is four units, so doubling both more than doubles the bill. Capacity added for a load test or a big reindex is easy to forget once the traffic settles down.
Checking query load against capacity
az search service list --resource-group <rg> \ --query "[].{name:name, sku:sku.name, replicas:replicaCount, partitions:partitionCount}" -o table
az monitor metrics list \ --resource /subscriptions/<sub>/resourceGroups/<rg>/providers/Microsoft.Search/searchServices/<service> \ --metric SearchQueriesPerSecond ThrottledSearchQueriesPercentage \ --aggregation Average --interval PT1H --offset 30dThe availability floor ZopNight respects
Microsoft’s reliability guidance
ties the SLA to replicas: at least two for read-only workloads and at least three for read-write
workloads. You can tell ZopNight which one you need with the tag zopnight:search-read-sla on the
service:
read(ortrue,1,yes) keeps 2 units.read-writekeeps 3 units.- No tag keeps the minimum of 1 replica by 1 partition.
Conditions for a scale-down finding
- The service has more units than its floor.
SearchQueriesPerSeconddata exists and its 30-day average is below 0.1.- That metric has at least 7 days of real coverage, so a service created last week is not judged.
ThrottledSearchQueriesPercentageis absent or averages below 1%. Any real throttling means the capacity is being used.- The service has a positive monthly price.
Services that are left as they are
A service already at its floor, including every single-unit service, gets no finding. Without a query metric, or with too little history, nothing is raised. The rule never guesses a saving: if the cost is unknown, there is no finding at all.
Splitting the bill by search unit
Every unit in a service bills at the same rate, so the per-unit cost is the service’s own cost divided by its unit count:
per-unit cost = monthly cost / (replicas x partitions)saving = per-unit cost x (current units - floor units), capped at monthly costScaling the service down
- Review query volume, latency and indexing schedules in the service’s Monitoring blade before changing anything.
- Reduce replicas, partitions or both:
az search service update --resource-group <rg> --name <service> --replica-count <n> --partition-count <n>. - Keep at least two replicas for a read SLA and three for a read-write SLA.