# Azure VM Without Availability Set

> Flags production Azure VMs with no availability set or zone, which can only be fixed by recreating the VM.

Source: https://zop.dev/integrations/azure/recommendations/azure-vm-without-availability-set

---

## A lone VM has no redundancy of its own

[Availability sets](https://learn.microsoft.com/en-us/azure/virtual-machines/availability-set-overview)
spread VMs across fault domains, which share power and network switches, and update domains,
which Azure restarts one at a time during planned maintenance. An availability set can have up to
3 fault domains and 20 update domains. Availability zones go further, placing VMs in physically
separate datacenters within a region.

A VM in neither has no placement guarantee relative to anything. Microsoft ties its 99.95% SLA to
two or more VMs in an availability set; a single VM outside one does not get that commitment.

## Listing VMs without placement

```bash
az vm list \
  --query "[?availabilitySet==null && zones==null].{name:name, rg:resourceGroup}" \
  -o table
```

Filter the output to production workloads; the same query returns dev and test VMs too.

## Gates the VM must pass

1. The VM is classified as production: its name contains `prod`, `production`, `prd` or `live`
   (and does not look like dev or test), or an `env`, `environment`, `stage` or `tier` tag says so.
2. It is not an instance of a Uniform orchestration scale set; those already get fault-domain and
   zone spread from the scale set.
3. Azure reports no availability set and no zone for the VM.

## Cases that do not fire, and one that might

Non-production VMs are not evaluated. If ZopNight could not determine the VM's placement, it
raises nothing. One known gap remains: members of Flexible orchestration scale sets look like
standalone VMs today, so a regional Flexible member can be flagged even though the scale set
handles its distribution. Dismiss those with that reason.

## Downtime risk, not a bill

There is no saving attached, and availability sets themselves cost nothing extra; you pay only
for the VMs. The risk is a production service going down with its only host during hardware
failure or maintenance.

## Rebuilding with redundancy

Microsoft states that [a VM can be added to an availability set only when it is created](https://learn.microsoft.com/en-us/azure/virtual-machines/windows/change-availability-set),
so the fix is a rebuild.

1. Decide between zones, for datacenter-level isolation, and an availability set, for lower
   VM-to-VM latency.
2. Set the VM's disks to detach rather than delete on VM deletion.
3. Delete the VM and recreate it from the same disks in the chosen set or zone.
4. Add at least a second VM behind a load balancer; one VM in a set gains little.
5. For fleets, consider a Virtual Machine Scale Set, which distributes instances automatically.

**Warning**
Deleting the VM with disks set to delete destroys the data. Check each disk's delete
option before you start.
