Skip to main content
resource · gcp

Vertex AI Dataset

schedulable
no
category
ai-ml-services

Does ZopNight manage Vertex AI Dataset?

Vertex AI datasets are lightweight registry objects: the managed dataset itself adds no meaningful charge, while the underlying training data bills through Cloud Storage or BigQuery. ZopNight inventories every dataset via Cloud Asset Inventory and records its modality: tabular, image, text, video, or time series.

Rules that fire on Vertex AI Dataset

no live rules

No active rule family targets Vertex AI Dataset today. Rules that used to are retired, and retired rules publish no pages and fire no findings. Scheduling and permissions coverage are unaffected.

Browse every live recommendation for this platform →

A Vertex AI dataset is a managed collection of training data for tabular, image, text, video, or time-series models. The dataset object itself is lightweight; the underlying data bills through Cloud Storage or BigQuery.

A pointer, not a payload

The managed dataset is essentially a structured reference: it records where training data lives and how it is labeled across the five supported modalities of tabular, image, text, video, and time series. The object adds no meaningful charge of its own. What bills is the substrate underneath: source files in Cloud Storage at that product’s storage rates, tabular sources under BigQuery storage meters, and any labeling work performed against the data.

What the dataset row records, and what it does not

ZopNight inventories datasets via Cloud Asset Inventory as standalone rows, each tagged with its modality. Discovery does not read training-input references, so ZopNight does not link a dataset to the pipelines that used it or to the bucket behind it. Answering “what data trained this model” still means opening the training pipeline’s input configuration in Vertex AI.

Training data that outlives its models

Dataset-shaped waste hides downstream of the object itself. Source buckets stay at full size for models that retrain quarterly from fresh data and will never reread the old corpus. Labeled datasets persist for projects that shipped or died. Duplicate exports of the same tables sit staged for experiments that ended without cleanup. ZopNight does not link datasets to pipelines or buckets, so finding these means opening each dataset’s source location in Vertex AI and checking whether any recent training run still reads it.

Managed datasets in the console

Google Cloud console → Vertex AI → Datasets lists datasets per region with modality and creation date. Cross-referencing that list against Model Registry activity is the quickest way to spot training data no pipeline has touched in months. Since the object itself is free, keeping datasets registered costs nothing. The audit value lies in the source locations each dataset records, which Vertex AI shows but ZopNight does not read.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·