# GPU Instance Low Utilization

> Finds GPU instances with low measured GPU utilization and prices one size down in the same GPU family.

Source: https://zop.dev/integrations/aws/recommendations/gpu-instance-low-utilization

---

## GPU hours are the most expensive hours in EC2

GPU instances cost far more per hour than general-purpose ones. The current On-Demand Linux rates
for US East (N. Virginia) put a `g5.xlarge` at $1.006 an hour, a `p3.2xlarge` at $3.06 and a
`p3.8xlarge` at $12.24. A `p3.8xlarge` left running for a month is close to $9,000, whether the GPUs
are training a model or waiting for the next job.

EC2 publishes no GPU metric by itself. The
[CloudWatch agent collects NVIDIA GPU metrics](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/CloudWatch-Agent-NVIDIA-GPU.html)
on Linux when you add an `nvidia_gpu` section to its configuration; `nvidia_smi_utilization_gpu` is
the percentage of time one or more kernels were running on the GPU.

## Reading GPU utilization

```bash
aws cloudwatch get-metric-statistics --namespace CWAgent \
  --metric-name nvidia_smi_utilization_gpu \
  --dimensions Name=InstanceId,Value=i-0123456789abcdef0 \
  --statistics Average Maximum --period 86400 \
  --start-time 2026-08-26T00:00:00Z --end-time 2026-09-25T00:00:00Z
```

The agent also tags each GPU metric with dimensions such as the GPU `index`, so list the metric with
`aws cloudwatch list-metrics --namespace CWAgent --metric-name nvidia_smi_utilization_gpu` first and
copy the exact dimension set.

## Conditions for a GPU downsize

1. The instance is running and belongs to a GPU family: `p3`, `p4`, `p4d`, `p5`, `g4`, `g4dn`,
   `g4ad`, `g5` or `g5g`.
2. ZopNight has the `nvidia_smi_utilization_gpu` series for it, spanning at least 7 days.
3. The 30-day average is under 20%, and where ZopNight holds at least 7 days of true hourly peaks,
   none reaches 80%.
4. A smaller size exists in the same family. If that size has fewer GPUs, the measured utilization
   must fit on them: moving from 4 GPUs to 2, for example, only goes ahead below 50%.
5. The instance's cost and both On-Demand rates are known.

## GPU boxes that are not flagged

An instance with no CloudWatch agent GPU data is skipped outright; there is no guessing from CPU.
Instances with under 7 days of history wait, so a box launched last week is not given a
high-severity call. General CPU rightsizing is done by
<a href="https://zop.dev/integrations/aws/recommendations/ec2-rightsizing">EC2 Rightsizing</a>, which leaves GPU
families to this rule.

## Two On-Demand rates, no assumed fraction

```text
saving = monthly cost x (1 - smaller size hourly rate / current hourly rate)
```

In us-east-1, `g5.2xlarge` to `g5.xlarge` goes from $1.212 to $1.006 an hour, about 17% of the
instance's cost. If either rate is missing, nothing is shown.

## Moving to the smaller GPU instance

1. Check GPU memory use as well as utilization; a smaller card can fail on memory before speed.
2. Validate the workload on the smaller type in staging.
3. Stop the instance, change its type, and start it:
   `aws ec2 modify-instance-attribute --instance-id i-0123456789abcdef0 --instance-type Value=g5.xlarge`
4. For work that runs in bursts, consider stopping the instance between jobs instead.

**Warning**
Changing the instance type requires a stop, which ends any job running on it.
