Vector Search index endpoints with near-zero CPU and no queries in 30 days
What does ZopNight detect here?
Vector Search charges per node hour for every VM hosting an index deployed to an endpoint, even when no query arrives. ZopNight flags an index endpoint whose `matching_engine/cpu/request_utilization` stayed under 1% on average and at peak with no recorded queries, across whatever history it holds, up to 42 days with no minimum, and recommends undeploying the index.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1217 |
| Category | idle |
| Severity | medium |
| Metric | matching_engine/cpu/request_utilization, matching_engine/query/request_count |
| Threshold | CPU request utilization under 1% and no queries |
| Evaluation window | up to 42d |
| Source | ZopNight |
| Permissions used | aiplatform.indexEndpoints.list · aiplatform.indexEndpoints.get · monitoring.timeSeries.list |
Where it applies
Deployed indexes keep serving nodes running
Vertex AI pricing splits Vector Search into two charges: building or updating indexes, billed per GiB processed, and serving, billed “per node hour” for each VM that hosts a deployed index. Google’s own estimate formula multiplies replicas, shards and the hourly node rate by 730 hours, so serving is the recurring part of the bill. The nodes exist because an index is deployed to an endpoint, not because queries are arriving.
Finding endpoints with no query traffic
List the index endpoints in a region, then describe any that look unfamiliar to see which deployed indexes they carry:
gcloud ai index-endpoints list --region=us-central1In Metrics Explorer, chart aiplatform.googleapis.com/matching_engine/cpu/request_utilization
and aiplatform.googleapis.com/matching_engine/query/request_count for the IndexEndpoint
resource over 30 days. The request count is labelled by deployed_index_id, so you can see which
deployed index, if any, is being queried.
The evidence this rule asks for
ZopNight anchors the decision on the endpoint’s CPU request utilization, a gauge that keeps reporting while an index is deployed even when no query arrives. That series must be present, and both its average and its peak must stay under 1% over whatever metric history ZopNight holds, up to 42 days. There is no minimum history, so an endpoint with only a few days of data can be flagged; the finding’s evidence shows the last 7 days.
The query request count is a delta metric and may have no points at all for a quiet endpoint, so a missing query series counts as consistent with idleness. A query series that exists and shows any request in the window is a veto: one query in the month keeps the endpoint off the list.
When a quiet endpoint is still left alone
No CPU utilization series means no finding, because a missing gauge says nothing about whether the endpoint is in use. Any recorded query also stops the rule. ZopNight also needs a billed cost for the endpoint; when its records attribute no cost to it, the rule stays silent instead of estimating a node-hour rate. You may see quiet endpoints in the console that this rule has not flagged for that reason; the commands above are the way to check them by hand.
Where the saving comes from
saving = billed monthly cost ZopNight holds for the endpointcost after fix = 0Undeploying the index
-
Confirm no service, pipeline or scheduled job queries the endpoint; search your code for the endpoint ID and the deployed index ID.
-
Undeploy it, following Google’s undeploy steps:
Terminal window gcloud ai index-endpoints undeploy-index INDEX_ENDPOINT_ID \--region=us-central1 --deployed-index-id=DEPLOYED_INDEX_ID -
If nothing else is deployed there, delete the endpoint with
gcloud ai index-endpoints delete INDEX_ENDPOINT_ID --region=us-central1. -
The index itself stays; if it is no longer needed either, see GCP Vertex AI Vector Search Index Not Deployed.