Skip to main content
discount · gcp

GKE node pools in non-production clusters still running standard VMs

resource types
1
rule IDs covered
1
severity
low

What does ZopNight detect here?

GKE node pools are flagged when the pool name, its cloud account name or an environment label marks it as dev, test, QA, staging, sandbox or demo and the pool is not already on Spot VMs. Pools created with `gcloud container node-pools create --spot` bill at Spot rates, and GKE reschedules Pods when a node is reclaimed.

Signal and threshold

How ZopNight evaluates GKE node pools in non-production clusters still running standard VMs.
Field Value
Rule IDsRC-114
Categorydiscount
Severitylow
Metricnone — pure configuration read
Thresholdnon-production node pool not on Spot VMs
SourceZopNight
Permissions usedcontainer.clusters.list · container.clusters.get

Why node pools suit Spot better than single VMs

A GKE node pool is a group of interchangeable VMs behind the Kubernetes scheduler. Google’s Spot VMs on GKE page describes what happens when Compute Engine reclaims one: GKE receives a preemption notice and, by default, gives Pods 15 seconds to shut down, followed by 15 seconds for critical system Pods, inside a 30-second graceful termination window. Controllers then recreate the evicted Pods on other nodes. Unlike preemptible VMs, Spot VMs have no 24-hour expiry.

That is why this rule skips the standby, stateful-service and Windows vetoes that GCP Spot VM Opportunity applies to single Compute Engine VMs. Spot nodes also carry the cloud.google.com/gke-spot=true label, which makes it easy to steer workloads on or off them.

Seeing which pools are on standard VMs

Terminal window
gcloud container node-pools list --cluster=CLUSTER_NAME --location=LOCATION \
--format="table(name,config.machineType,config.spot)"

A pool with config.spot empty or false is on standard VMs.

What marks a pool as a candidate

ZopNight flags a node pool when all of these hold:

  1. No environment label on the pool says production.
  2. The pool name, the name of the cloud account it belongs to, or an environment label contains dev, test, qa, staging, sandbox or demo.
  3. The pool is not already recorded as Spot.
  4. The pool has a known monthly cost.

Pools that never get this finding

A production environment label vetoes the recommendation regardless of names. Pools already on Spot are skipped, and so is any pool whose cost ZopNight cannot price. The rule does not inspect the workloads on the pool, so check for stateful or long-running jobs yourself.

Estimating the Spot saving

Terminal window
saving = pool monthly cost x (1 - Spot rate / on-demand rate)

The discount comes from Google’s live on-demand and Spot rates for the pool’s machine type. When live Spot rates are missing, ZopNight uses a published baseline Spot discount for GCP instead, still applied to the pool’s real cost.

Moving workloads to a Spot pool

  1. Create a matching Spot pool: gcloud container node-pools create POOL_NAME-spot --cluster=CLUSTER_NAME --location=LOCATION --machine-type=MACHINE_TYPE --spot.
  2. Keep at least one standard pool so critical Pods, and Jobs when Spot capacity runs out, have somewhere to land, as Google’s best practices advise.
  3. Cordon and drain the old pool: kubectl cordon then kubectl drain --ignore-daemonsets for each node.
  4. Delete the old pool once Pods are running on the Spot nodes.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·