Skip to main content
idle · gcp

Vector Search index endpoints with near-zero CPU and no queries in 30 days

resource types
1
rule IDs covered
1
severity
medium

What does ZopNight detect here?

Vector Search charges per node hour for every VM hosting an index deployed to an endpoint, even when no query arrives. ZopNight flags an index endpoint whose `matching_engine/cpu/request_utilization` stayed under 1% on average and at peak with no recorded queries, across whatever history it holds, up to 42 days with no minimum, and recommends undeploying the index.

Signal and threshold

How ZopNight evaluates Vector Search index endpoints with near-zero CPU and no queries in 30 days.
Field Value
Rule IDsRC-1217
Categoryidle
Severitymedium
Metricmatching_engine/cpu/request_utilization, matching_engine/query/request_count
ThresholdCPU request utilization under 1% and no queries
Evaluation windowup to 42d
SourceZopNight
Permissions usedaiplatform.indexEndpoints.list · aiplatform.indexEndpoints.get · monitoring.timeSeries.list

Deployed indexes keep serving nodes running

Vertex AI pricing splits Vector Search into two charges: building or updating indexes, billed per GiB processed, and serving, billed “per node hour” for each VM that hosts a deployed index. Google’s own estimate formula multiplies replicas, shards and the hourly node rate by 730 hours, so serving is the recurring part of the bill. The nodes exist because an index is deployed to an endpoint, not because queries are arriving.

Finding endpoints with no query traffic

List the index endpoints in a region, then describe any that look unfamiliar to see which deployed indexes they carry:

Terminal window
gcloud ai index-endpoints list --region=us-central1

In Metrics Explorer, chart aiplatform.googleapis.com/matching_engine/cpu/request_utilization and aiplatform.googleapis.com/matching_engine/query/request_count for the IndexEndpoint resource over 30 days. The request count is labelled by deployed_index_id, so you can see which deployed index, if any, is being queried.

The evidence this rule asks for

ZopNight anchors the decision on the endpoint’s CPU request utilization, a gauge that keeps reporting while an index is deployed even when no query arrives. That series must be present, and both its average and its peak must stay under 1% over whatever metric history ZopNight holds, up to 42 days. There is no minimum history, so an endpoint with only a few days of data can be flagged; the finding’s evidence shows the last 7 days.

The query request count is a delta metric and may have no points at all for a quiet endpoint, so a missing query series counts as consistent with idleness. A query series that exists and shows any request in the window is a veto: one query in the month keeps the endpoint off the list.

When a quiet endpoint is still left alone

No CPU utilization series means no finding, because a missing gauge says nothing about whether the endpoint is in use. Any recorded query also stops the rule. ZopNight also needs a billed cost for the endpoint; when its records attribute no cost to it, the rule stays silent instead of estimating a node-hour rate. You may see quiet endpoints in the console that this rule has not flagged for that reason; the commands above are the way to check them by hand.

Where the saving comes from

Terminal window
saving = billed monthly cost ZopNight holds for the endpoint
cost after fix = 0

Undeploying the index

  1. Confirm no service, pipeline or scheduled job queries the endpoint; search your code for the endpoint ID and the deployed index ID.

  2. Undeploy it, following Google’s undeploy steps:

    Terminal window
    gcloud ai index-endpoints undeploy-index INDEX_ENDPOINT_ID \
    --region=us-central1 --deployed-index-id=DEPLOYED_INDEX_ID
  3. If nothing else is deployed there, delete the endpoint with gcloud ai index-endpoints delete INDEX_ENDPOINT_ID --region=us-central1.

  4. The index itself stays; if it is no longer needed either, see GCP Vertex AI Vector Search Index Not Deployed.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·