Skip to main content
rightsizing · aws

Amazon ECS services averaging under 10% CPU of their task reservation over 30 days

resource types
1
rule IDs covered
1
severity
medium

What does ZopNight detect here?

ZopNight flags Amazon ECS services whose service-level `CPUUtilization` averages under 10% over 30 days, with memory under 50% when measured. It recommends the next smaller task CPU size, keeps memory in a valid pairing, and prices the saving from Fargate per-task rates; the resize can be applied as a rolling service update.

Signal and threshold

How ZopNight evaluates Amazon ECS services averaging under 10% CPU of their task reservation over 30 days.
Field Value
Rule IDsRC-158
Categoryrightsizing
Severitymedium
MetricCPUUtilization
Threshold< 10% average CPU, < 50% memory if measured
Evaluation window30d
SourceZopNight
Permissions usedecs:ListServices · ecs:DescribeServices · ecs:DescribeTaskDefinition · cloudwatch:GetMetricStatistics

Service CPU is measured against what the tasks reserve

For an ECS service, AWS defines CPUUtilization as the CPU units in use by the service’s tasks divided by the CPU units reserved for them. A value of 6% means the tasks use about one sixteenth of the CPU they were sized for. On Fargate that reservation is exactly what you are billed for, so a low figure is a direct measure of overpayment.

ECS publishes these metrics every minute, and only for services with tasks in the RUNNING state.

Reading utilisation for a service

Terminal window
aws ecs list-services --cluster prod
aws cloudwatch get-metric-statistics --namespace AWS/ECS --metric-name CPUUtilization \
--dimensions Name=ClusterName,Value=prod Name=ServiceName,Value=api \
--statistics Average Maximum --period 86400 \
--start-time 2026-08-26T00:00:00Z --end-time 2026-09-25T00:00:00Z

Run the same query for MemoryUtilization, then read the task size from the service’s current task definition with aws ecs describe-task-definition.

Thresholds and guards on each service

  1. The service is active, not draining or stopped, and has desired and running tasks above zero.
  2. The 30-day CPU average is under 10%, and the CPU series does not show a peak at 80% or more.
  3. If memory is measured, its 30-day average is under 50%. A service with no memory series is still evaluated.
  4. The task definition’s CPU and memory are known, and the CPU value is on the standard Fargate size ladder with a smaller size below it.
  5. A monthly cost is known for the service.

Services that are not resized

Services on the smallest CPU size, 256 units, have no step down. Services with an off-ladder CPU value, or whose memory cannot be paired with the smaller CPU size, are skipped rather than guessed. Services running nothing are left to ECS Idle Service. Very large tasks above 4 vCPU or 8 GiB are also checked by ECS Fargate Oversized Task; if both land on the same saving, this recommendation is the one kept.

Why the saving uses per-task rates, not the CPU ratio

Terminal window
saving = monthly cost x (1 - rate of smaller task size / rate of current task size)

Only the vCPU part of the price shrinks when memory stays the same, so scaling the whole bill by the CPU ratio would overstate the saving. The ratio of the two task rates avoids that.

Applying the smaller size

ZopNight can apply this as a guided change: it registers a task definition revision with the smaller CPU value and updates the service to it. To do it by hand:

  1. Copy the current task definition, lower cpu by one size and adjust memory only if the new CPU size requires it.
  2. Register it with aws ecs register-task-definition --cli-input-json file://api.json.
  3. Roll the service onto it with aws ecs update-service --cluster prod --service api --task-definition api:43.
  4. Check CPU, latency and error rates after the deployment completes.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·