Skip to main content
idle · gcp

Vertex AI endpoints with deployed models that served zero predictions

resource types
1
rule IDs covered
1
severity
medium

What does ZopNight detect here?

Vertex AI endpoints are flagged when `aiplatform.googleapis.com/prediction/online/prediction_count` shows zero predictions, on average and at peak, across the 30-day lookback. Google bills every model deployed to an endpoint per node hour even if no prediction is made, and charges stop only when the model is undeployed.

Signal and threshold

How ZopNight evaluates Vertex AI endpoints with deployed models that served zero predictions.
Field Value
Rule IDsRC-1212
Categoryidle
Severitymedium
Metricaiplatform.googleapis.com/prediction/online/prediction_count
Thresholdzero predictions, average and peak
Evaluation window30d
SourceZopNight
Permissions usedaiplatform.endpoints.list · aiplatform.endpoints.get · monitoring.timeSeries.list

Deployed models bill while they wait

Online prediction on Vertex AI is priced in node hours, and Google’s Vertex AI pricing defines a node hour as time a VM spends running a prediction job or waiting in an active state, meaning an endpoint with one or more models deployed, ready to answer. The same page is blunt about the consequence: you pay for each model deployed to an endpoint even if no prediction is made, and you must undeploy the model to stop further charges.

So an endpoint that nobody calls is not free. Its serving nodes are allocated and billed so that they can respond immediately, whatever the request rate.

Checking an endpoint’s traffic

Terminal window
gcloud ai endpoints list --region=REGION
gcloud ai endpoints describe ENDPOINT_ID --region=REGION

The describe output lists each deployed model and its ID. For traffic, chart aiplatform.googleapis.com/prediction/online/prediction_count for the endpoint in Metrics Explorer over 30 days.

When a finding is created

ZopNight reads the endpoint’s prediction count by name and needs both the average and the peak to be zero across the data it collected. The endpoint must also carry a real cost in your billing data; ZopNight does not estimate serving rates for Vertex AI endpoints.

When ZopNight also has hour-by-hour usage patterns for the endpoint, it recommends a recurring off-hours window instead of deletion: undeploy the model for the measured idle hours and redeploy before working hours, which is reversible and keeps daytime serving unchanged.

When it holds back

No prediction data means no finding. A single prediction in the window clears the endpoint. Endpoints with no attributed cost are skipped rather than priced from list rates. Undeployed models and other Vertex AI resources have their own checks, such as GCP Vertex AI Model Not Deployed.

Two ways the saving is counted

Terminal window
idle teardown: saving = current monthly endpoint cost, cost after = 0
off-hours window: saving = monthly cost x measured idle share of the week

Undeploying the model

  1. Confirm with the owners that no application, batch job or evaluation calls the endpoint.
  2. Find the deployed model ID with gcloud ai endpoints describe.
  3. Undeploy it: gcloud ai endpoints undeploy-model ENDPOINT_ID --region=REGION --deployed-model-id=DEPLOYED_MODEL_ID.
  4. Optionally delete the endpoint with gcloud ai endpoints delete ENDPOINT_ID --region=REGION if nothing will be deployed to it again.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·