Vertex AI endpoints with deployed models that served zero predictions
What does ZopNight detect here?
Vertex AI endpoints are flagged when `aiplatform.googleapis.com/prediction/online/prediction_count` shows zero predictions, on average and at peak, across the 30-day lookback. Google bills every model deployed to an endpoint per node hour even if no prediction is made, and charges stop only when the model is undeployed.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1212 |
| Category | idle |
| Severity | medium |
| Metric | aiplatform.googleapis.com/prediction/online/prediction_count |
| Threshold | zero predictions, average and peak |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | aiplatform.endpoints.list · aiplatform.endpoints.get · monitoring.timeSeries.list |
Where it applies
Deployed models bill while they wait
Online prediction on Vertex AI is priced in node hours, and Google’s Vertex AI pricing defines a node hour as time a VM spends running a prediction job or waiting in an active state, meaning an endpoint with one or more models deployed, ready to answer. The same page is blunt about the consequence: you pay for each model deployed to an endpoint even if no prediction is made, and you must undeploy the model to stop further charges.
So an endpoint that nobody calls is not free. Its serving nodes are allocated and billed so that they can respond immediately, whatever the request rate.
Checking an endpoint’s traffic
gcloud ai endpoints list --region=REGIONgcloud ai endpoints describe ENDPOINT_ID --region=REGIONThe describe output lists each deployed model and its ID. For traffic, chart
aiplatform.googleapis.com/prediction/online/prediction_count for the endpoint in Metrics
Explorer over 30 days.
When a finding is created
ZopNight reads the endpoint’s prediction count by name and needs both the average and the peak to be zero across the data it collected. The endpoint must also carry a real cost in your billing data; ZopNight does not estimate serving rates for Vertex AI endpoints.
When ZopNight also has hour-by-hour usage patterns for the endpoint, it recommends a recurring off-hours window instead of deletion: undeploy the model for the measured idle hours and redeploy before working hours, which is reversible and keeps daytime serving unchanged.
When it holds back
No prediction data means no finding. A single prediction in the window clears the endpoint. Endpoints with no attributed cost are skipped rather than priced from list rates. Undeployed models and other Vertex AI resources have their own checks, such as GCP Vertex AI Model Not Deployed.
Two ways the saving is counted
idle teardown: saving = current monthly endpoint cost, cost after = 0off-hours window: saving = monthly cost x measured idle share of the weekUndeploying the model
- Confirm with the owners that no application, batch job or evaluation calls the endpoint.
- Find the deployed model ID with
gcloud ai endpoints describe. - Undeploy it:
gcloud ai endpoints undeploy-model ENDPOINT_ID --region=REGION --deployed-model-id=DEPLOYED_MODEL_ID. - Optionally delete the endpoint with
gcloud ai endpoints delete ENDPOINT_ID --region=REGIONif nothing will be deployed to it again.