Skip to main content
rightsizing · aws

Running EC2 GPU instances averaging under 20% GPU utilization over 30 days

resource types
1
rule IDs covered
1
severity
high

What does ZopNight detect here?

ZopNight flags running EC2 GPU instances in families such as `p3`, `p4d`, `p5`, `g4dn` and `g5` whose `nvidia_smi_utilization_gpu` from the CloudWatch agent averages under 20% over 30 days, with at least 7 days of data. It recommends the next smaller size in the same GPU family and prices the saving from the two On-Demand rates.

Signal and threshold

How ZopNight evaluates Running EC2 GPU instances averaging under 20% GPU utilization over 30 days.
Field Value
Rule IDsRC-1500
Categoryrightsizing
Severityhigh
Metricnvidia_smi_utilization_gpu
Threshold< 20% average
Evaluation window30d
SourceZopNight
Permissions usedec2:DescribeInstances · cloudwatch:GetMetricStatistics

GPU hours are the most expensive hours in EC2

GPU instances cost far more per hour than general-purpose ones. The current On-Demand Linux rates for US East (N. Virginia) put a g5.xlarge at $1.006 an hour, a p3.2xlarge at $3.06 and a p3.8xlarge at $12.24. A p3.8xlarge left running for a month is close to $9,000, whether the GPUs are training a model or waiting for the next job.

EC2 publishes no GPU metric by itself. The CloudWatch agent collects NVIDIA GPU metrics on Linux when you add an nvidia_gpu section to its configuration; nvidia_smi_utilization_gpu is the percentage of time one or more kernels were running on the GPU.

Reading GPU utilization

Terminal window
aws cloudwatch get-metric-statistics --namespace CWAgent \
--metric-name nvidia_smi_utilization_gpu \
--dimensions Name=InstanceId,Value=i-0123456789abcdef0 \
--statistics Average Maximum --period 86400 \
--start-time 2026-08-26T00:00:00Z --end-time 2026-09-25T00:00:00Z

The agent also tags each GPU metric with dimensions such as the GPU index, so list the metric with aws cloudwatch list-metrics --namespace CWAgent --metric-name nvidia_smi_utilization_gpu first and copy the exact dimension set.

Conditions for a GPU downsize

  1. The instance is running and belongs to a GPU family: p3, p4, p4d, p5, g4, g4dn, g4ad, g5 or g5g.
  2. ZopNight has the nvidia_smi_utilization_gpu series for it, spanning at least 7 days.
  3. The 30-day average is under 20%, and where ZopNight holds at least 7 days of true hourly peaks, none reaches 80%.
  4. A smaller size exists in the same family. If that size has fewer GPUs, the measured utilization must fit on them: moving from 4 GPUs to 2, for example, only goes ahead below 50%.
  5. The instance’s cost and both On-Demand rates are known.

GPU boxes that are not flagged

An instance with no CloudWatch agent GPU data is skipped outright; there is no guessing from CPU. Instances with under 7 days of history wait, so a box launched last week is not given a high-severity call. General CPU rightsizing is done by EC2 Rightsizing, which leaves GPU families to this rule.

Two On-Demand rates, no assumed fraction

Terminal window
saving = monthly cost x (1 - smaller size hourly rate / current hourly rate)

In us-east-1, g5.2xlarge to g5.xlarge goes from $1.212 to $1.006 an hour, about 17% of the instance’s cost. If either rate is missing, nothing is shown.

Moving to the smaller GPU instance

  1. Check GPU memory use as well as utilization; a smaller card can fail on memory before speed.
  2. Validate the workload on the smaller type in staging.
  3. Stop the instance, change its type, and start it: aws ec2 modify-instance-attribute --instance-id i-0123456789abcdef0 --instance-type Value=g5.xlarge
  4. For work that runs in bursts, consider stopping the instance between jobs instead.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·