Skip to main content
discount · azure

Dev and test AKS node pools running on regular-priority VMs instead of Spot

resource types
1
rule IDs covered
1
severity
low

What does ZopNight detect here?

ZopNight flags an AKS node pool whose name or environment tag marks it as dev, test, QA, staging, sandbox or demo but whose nodes run at regular priority. A secondary pool with `--priority Spot` would run the same interruptible workloads on evictable capacity, and ZopNight prices the gap from the live pay-as-you-go and Spot rates for that VM size.

Signal and threshold

How ZopNight evaluates Dev and test AKS node pools running on regular-priority VMs instead of Spot.
Field Value
Rule IDsRC-250
Categorydiscount
Severitylow
Metricnone — pure configuration read
Thresholddev/test name or env tag, Spot not enabled
SourceZopNight
Permissions usedMicrosoft.ContainerService/managedClusters/read · Microsoft.ContainerService/managedClusters/agentPools/read

What a Spot node pool changes on the bill

AKS node pools are virtual machine scale sets, and each VM is billed at its size’s rate whether the pods on it are busy or not. A Spot node pool is backed by an Azure Spot scale set instead, which uses spare Azure capacity at a discount. The trade is stated plainly by Microsoft: there is no SLA and no high-availability guarantee for Spot nodes, and Azure evicts them when it needs the capacity back. Microsoft lists development and testing environments and batch jobs as the workloads that fit.

A dev or staging cluster is usually the easiest place to take that trade, because an evicted node there costs a pod restart, not an outage.

Checking node pool priority on a cluster

Terminal window
az aks nodepool list --resource-group my-rg --cluster-name my-aks \
--query "[].{name:name, mode:mode, priority:scaleSetPriority, size:vmSize, count:count}" \
-o table

A pool showing Regular (or an empty priority) in a non-production cluster is a candidate. Spot pools show Spot and carry the kubernetes.azure.com/scalesetpriority=spot:NoSchedule taint.

Signals ZopNight needs before it raises the finding

  1. A positive non-production signal: an environment tag with a dev or test value, or a pool name containing dev, test, qa, staging, sandbox or demo. Names that only look like dev, such as devops, runner, agent, bastion or jumpbox, do not count.
  2. No production environment tag. A prod, production, prd or live value on an env, environment, stage or tier tag overrides any dev-looking name.
  3. Spot is not already on: the pool’s Spot flag is not true and, when that flag is missing, it has no spot=true tag.
  4. A known monthly cost for the pool, plus pay-as-you-go and Spot rates for its VM size.

Pools that get no Spot suggestion

A pool without pricing for both rate tiers gets no finding at all. ZopNight does not fall back to a typical Spot discount, because Spot prices vary by size, region and time. A computed discount outside a plausible 5% to 92% band is treated as bad rate data and also produces nothing. Pools with a neutral name and no environment tag are left alone even if they are in fact test pools: the rule needs a positive signal, not the absence of a production one.

Spot rate against the regular rate, applied to the pool

Terminal window
discount = 1 - (Spot hourly rate / pay-as-you-go hourly rate) for the pool's VM size
saving = current monthly pool cost x discount

The figure assumes the workload moves to Spot nodes of the same size. It is a ceiling: pods that must stay on regular nodes, and the capacity you keep for them, reduce it.

Moving interruptible workloads to a Spot pool

A pool’s priority cannot be changed after creation, so the move is additive:

  1. Add a secondary Spot pool: az aks nodepool add --resource-group my-rg --cluster-name my-aks --name spotpool --priority Spot --eviction-policy Delete --spot-max-price -1 --enable-cluster-autoscaler --min-count 1 --max-count 3. A Spot pool cannot be the cluster’s default pool.
  2. Give the workloads that may move a toleration for the Spot taint and a node affinity for the kubernetes.azure.com/scalesetpriority: spot label.
  3. Keep the cluster autoscaler on, so evicted capacity is replaced when Spot is available again.
  4. Scale the original regular pool down to what must stay on regular nodes, and watch pod evictions for a week before shrinking further.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·