Dev and test AKS node pools running on regular-priority VMs instead of Spot
What does ZopNight detect here?
ZopNight flags an AKS node pool whose name or environment tag marks it as dev, test, QA, staging, sandbox or demo but whose nodes run at regular priority. A secondary pool with `--priority Spot` would run the same interruptible workloads on evictable capacity, and ZopNight prices the gap from the live pay-as-you-go and Spot rates for that VM size.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-250 |
| Category | discount |
| Severity | low |
| Metric | none — pure configuration read |
| Threshold | dev/test name or env tag, Spot not enabled |
| Source | ZopNight |
| Permissions used | Microsoft.ContainerService/managedClusters/read · Microsoft.ContainerService/managedClusters/agentPools/read |
Where it applies
What a Spot node pool changes on the bill
AKS node pools are virtual machine scale sets, and each VM is billed at its size’s rate whether the pods on it are busy or not. A Spot node pool is backed by an Azure Spot scale set instead, which uses spare Azure capacity at a discount. The trade is stated plainly by Microsoft: there is no SLA and no high-availability guarantee for Spot nodes, and Azure evicts them when it needs the capacity back. Microsoft lists development and testing environments and batch jobs as the workloads that fit.
A dev or staging cluster is usually the easiest place to take that trade, because an evicted node there costs a pod restart, not an outage.
Checking node pool priority on a cluster
az aks nodepool list --resource-group my-rg --cluster-name my-aks \ --query "[].{name:name, mode:mode, priority:scaleSetPriority, size:vmSize, count:count}" \ -o tableA pool showing Regular (or an empty priority) in a non-production cluster is a candidate.
Spot pools show Spot and carry the kubernetes.azure.com/scalesetpriority=spot:NoSchedule taint.
Signals ZopNight needs before it raises the finding
- A positive non-production signal: an environment tag with a dev or test value, or a pool
name containing
dev,test,qa,staging,sandboxordemo. Names that only look like dev, such asdevops,runner,agent,bastionorjumpbox, do not count. - No production environment tag. A
prod,production,prdorlivevalue on anenv,environment,stageortiertag overrides any dev-looking name. - Spot is not already on: the pool’s Spot flag is not
trueand, when that flag is missing, it has nospot=truetag. - A known monthly cost for the pool, plus pay-as-you-go and Spot rates for its VM size.
Pools that get no Spot suggestion
A pool without pricing for both rate tiers gets no finding at all. ZopNight does not fall back to a typical Spot discount, because Spot prices vary by size, region and time. A computed discount outside a plausible 5% to 92% band is treated as bad rate data and also produces nothing. Pools with a neutral name and no environment tag are left alone even if they are in fact test pools: the rule needs a positive signal, not the absence of a production one.
Spot rate against the regular rate, applied to the pool
discount = 1 - (Spot hourly rate / pay-as-you-go hourly rate) for the pool's VM sizesaving = current monthly pool cost x discountThe figure assumes the workload moves to Spot nodes of the same size. It is a ceiling: pods that must stay on regular nodes, and the capacity you keep for them, reduce it.
Moving interruptible workloads to a Spot pool
A pool’s priority cannot be changed after creation, so the move is additive:
- Add a secondary Spot pool:
az aks nodepool add --resource-group my-rg --cluster-name my-aks --name spotpool --priority Spot --eviction-policy Delete --spot-max-price -1 --enable-cluster-autoscaler --min-count 1 --max-count 3. A Spot pool cannot be the cluster’s default pool. - Give the workloads that may move a toleration for the Spot taint and a node affinity for the
kubernetes.azure.com/scalesetpriority: spotlabel. - Keep the cluster autoscaler on, so evicted capacity is replaced when Spot is available again.
- Scale the original regular pool down to what must stay on regular nodes, and watch pod evictions for a week before shrinking further.