Skip to main content
resource · gcp

Vertex AI Pipeline Job

schedulable
no
category
ai-ml-services

Does ZopNight manage Vertex AI Pipeline Job?

Vertex AI Pipelines charge a small per-run orchestration fee plus the compute every step consumes, and step compute dominates the total. ZopNight discovers pipeline jobs through the live aiplatform API rather than stale CAI rows and attributes per-step spend, surfacing failed pipelines that re-run expensive training steps.

Rules that fire on Vertex AI Pipeline Job

no live rules

No active rule family targets Vertex AI Pipeline Job today. Rules that used to are retired, and retired rules publish no pages and fire no findings. Scheduling and permissions coverage are unaffected.

Browse every live recommendation for this platform →

A pipeline job orchestrates a DAG of ML steps on Vertex AI Pipelines, billed a small per-run fee plus the compute each step consumes. Failed pipelines that re-run expensive training steps are the main waste pattern.

A small orchestration fee plus every step’s compute

Vertex AI Pipelines charges a flat fee per pipeline run, and then each step bills for whatever compute it uses: training jobs, batch predictions, data processing. Step compute dominates; the orchestration fee is minor. That structure means the pipeline object itself is cheap while its contents can be arbitrarily expensive, and the true total only appears when per-step spend is summed across the whole DAG.

Pipeline-job visibility in ZopNight

ZopNight discovers pipeline jobs through the live aiplatform API across 32 Vertex regions rather than relying on Cloud Asset Inventory, whose rows for transient jobs go stale after completion. Per-step compute spend is attributed in ML cost views, so a pipeline’s real cost (the sum of its steps) is visible in one place instead of scattered across anonymous training and prediction charges. Pipeline jobs are not schedulable; each run terminates by itself.

Re-run waste in ML pipelines

The dominant leak is the failed run repeating its expensive early steps: a pipeline that trains for hours and fails at a final deployment step will, without step caching, re-train from scratch on retry. Add cron-scheduled pipelines still executing after their consumers disappeared, and individual steps specified with accelerators sized for a different workload entirely. Reviewing where the DAG fails is worth far more than reviewing the orchestration fee.

Finding pipeline runs in the console

Google Cloud console → Vertex AI → Pipelines lists runs per region with state, duration, and a step-by-step breakdown. That step view is exactly where a failed run’s repeated cost shows up, and where step caching proves whether a retry actually skipped the expensive work.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

417 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

417 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·