Skip to main content
rightsizing · azure

Azure OpenAI provisioned deployments using less than half of their PTUs

resource types
1
rule IDs covered
1
severity
medium

What does ZopNight detect here?

ZopNight flags an Azure OpenAI account whose `AzureOpenAIProvisionedManagedUtilizationV2` metric averages below 50% over 30 days, with at least 7 days of data. It sizes the provisioned deployments down to observed demand, respects each deployment type's minimum and scale increment, and prices the removed PTUs at the real hourly PTU rate.

Signal and threshold

How ZopNight evaluates Azure OpenAI provisioned deployments using less than half of their PTUs.
Field Value
Rule IDsRC-1395
Categoryrightsizing
Severitymedium
MetricProvisioned-managed Utilization V2
Thresholdaverage utilization < 50%
Evaluation window30d
SourceZopNight
Permissions usedMicrosoft.CognitiveServices/accounts/read · Microsoft.CognitiveServices/accounts/deployments/read · Microsoft.Insights/Metrics/Read

Paying for PTUs by the hour whether or not tokens flow

Provisioned throughput is billed on capacity, not consumption. The provisioned throughput billing guide says provisioned deployments (Global, Data Zone and Regional) are charged an hourly rate per PTU for every PTU deployed, and that you are billed for the full deployed count regardless of actual utilization. Azure Reservations lower that hourly rate in exchange for a one-month or one-year commitment, but do not change the fact that the capacity is paid for in full.

A deployment sized for a launch forecast that never arrived therefore bills for its idle PTUs every hour. The utilization metric is the only place that shows it.

Reading deployed PTUs and utilization

List the deployments on an account with their SKU and PTU count:

Terminal window
az cognitiveservices account deployment list --resource-group my-rg --name my-openai \
--query "[].{name:name, sku:sku.name, ptu:sku.capacity}" -o table

Then read 30 days of utilization for the account. The metric is only emitted by provisioned deployments:

Terminal window
az monitor metrics list --resource <openai-account-resource-id> \
--metric AzureOpenAIProvisionedManagedUtilizationV2 \
--aggregation Average Maximum --interval PT24H --offset 30d

Utilization and coverage conditions

  1. The account reports Provisioned-managed Utilization V2, which means it runs at least one provisioned deployment. Standard, pay-as-you-go deployments never emit it and are never flagged.
  2. The metric covers at least 7 days, so a freshly deployed account warming up is not judged on a few quiet hours.
  3. Average utilization over 30 days is below 50%.
  4. Every provisioned deployment on the account (GlobalProvisionedManaged, DataZoneProvisionedManaged or ProvisionedManaged) has a known PTU count and a known hourly PTU rate for its type.
  5. The right-sized PTU count is actually lower than what is deployed.

Accounts that get no recommendation

If any provisioned deployment’s PTU count is unknown, the whole account is skipped rather than under-counted. No rate for the region, thin metric coverage, utilization at or above 50%, or a deployment already at its minimum size also mean no finding. One known limitation: the metric is read at account level, so an account with several provisioned deployments is judged on their blended average rather than each deployment separately.

PTU arithmetic behind the saving

Terminal window
current monthly = sum over deployments of (PTUs x hourly PTU rate x 730)
right-sized PTUs = deployed PTUs x utilization, rounded up to the type's scale increment,
never below the type's minimum deployment size
saving = (current monthly / deployed PTUs) x (deployed PTUs - right-sized PTUs)

Each deployment is priced at its own type’s rate, so an account mixing Global and Regional deployments is priced correctly. When types are mixed, the minimum and increment used are the conservative combination of each deployment’s own values.

Scaling a provisioned deployment down

  1. In the Foundry portal, open the account’s Deployments and select the provisioned deployment.
  2. Review the Provisioned-managed Utilization V2 trend, paying attention to peaks.
  3. Reduce the PTU count toward the target in the recommendation, or recreate the workload as a Standard deployment if demand is small and spiky.
  4. If the PTUs are covered by a reservation, right-size at renewal. Microsoft notes that a reservation keeps its original quantity when you scale a deployment down, so reducing mid-term leaves unused reservation coverage rather than saving money.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·