Skip to main content
resource · gcp

Vertex AI Deployment Resource Pool

schedulable
no
category
ai-ml-services

Does ZopNight manage Vertex AI Deployment Resource Pool?

A Vertex AI deployment resource pool bills continuously for its shared serving VMs, no matter how many of the models packed onto it still receive traffic. ZopNight lists pools from the live aiplatform API across 30+ regions, since Cloud Asset Inventory does not index the type, and flags pools whose models receive no traffic.

Rules that fire on Vertex AI Deployment Resource Pool

no live rules

No active rule family targets Vertex AI Deployment Resource Pool today. Rules that used to are retired, and retired rules publish no pages and fire no findings. Scheduling and permissions coverage are unaffected.

Browse every live recommendation for this platform →

A deployment resource pool lets multiple models share the same provisioned serving VMs. The pool bills for its machines continuously, so pools kept warm for retired models waste serving capacity.

One set of VMs, many models, one continuous bill

A deployment resource pool inverts the usual Vertex serving arithmetic: instead of each deployed model holding its own serving nodes, several models share one provisioned pool of machines. The pool’s VMs bill for every hour they stand, regardless of how many of the tenant models still receive traffic. Sharing lowers cost per model while the tenants are alive; once models retire, the same sharing hides the waste, because no single idle model looks responsible for the still-running pool.

Live listing where Cloud Asset Inventory has no row

Cloud Asset Inventory does not index this type at all, so ZopNight lists deployment resource pools from the live aiplatform REST API region by region (sweeping every documented GA Vertex location, more than 30 of them) and maps each result through the same enrichment path used for asset-inventory resources. The inventory therefore reflects live state rather than a stale snapshot, and pools whose models receive no traffic are flagged.

Warm pools serving retired tenants

Waste concentrates in two shapes. A pool created for a family of experiment models keeps running after every tenant was undeployed except one low-value straggler. And a pool sized for the peak of a shared workload keeps that machine count while the surviving models would fit in a fraction of it. Because the pool bills as a unit, the remedy is consolidation: undeploy the stragglers, migrate anything still needed, then shrink or delete the pool.

Where resource pools surface

Deployment resource pools are primarily an API-level object. The console’s Vertex AI → Online prediction pages show the endpoints and deployed models that share a pool, which is where the tenancy, and the case for retiring the pool, becomes visible.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

417 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

417 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·