Skip to main content
rightsizing · aws

EKS managed node groups above their minimum size with nodes under 40% CPU and 50% memory for 14 days

resource types
1
rule IDs covered
1
severity
medium

What does ZopNight detect here?

ZopNight flags on-demand EKS managed node groups running more nodes than their minimum size when Container Insights shows average `node_cpu_utilization` at or under 40% and `node_memory_utilization` at or under 50% over 14 days. It recommends removing one node and prices the saving at that instance type's On-Demand rate for 730 hours.

Signal and threshold

How ZopNight evaluates EKS managed node groups above their minimum size with nodes under 40% CPU and 50% memory for 14 days.
Field Value
Rule IDsRC-042
Categoryrightsizing
Severitymedium
Metricnode_cpu_utilization, node_memory_utilization
ThresholdCPU <= 40% and memory <= 50%, desired > min
Evaluation window14d
SourceZopNight
Permissions usedeks:ListNodegroups · eks:DescribeNodegroup · cloudwatch:GetMetricStatistics

Each extra worker node is a full EC2 bill

A managed node group is an EC2 Auto Scaling group that Amazon EKS manages for you, and every node in it is an EC2 instance billed at its normal rate. The group’s desired size is the number of those instances running right now. If the scheduler has room to spare on every node, the last node added is paying for nothing but headroom.

Kubernetes requests decide where pods land, but the bill follows the node count. A cluster whose nodes average a third of their CPU and less than half their memory can usually lose a node without any pod going unscheduled.

Checking node-group size and node load

Terminal window
aws eks describe-nodegroup --cluster-name prod --nodegroup-name general \
--query 'nodegroup.[status,capacityType,instanceTypes,scalingConfig]'
aws cloudwatch get-metric-statistics --namespace ContainerInsights \
--metric-name node_cpu_utilization --dimensions Name=ClusterName,Value=prod \
--statistics Average --period 86400 \
--start-time 2026-09-11T00:00:00Z --end-time 2026-09-25T00:00:00Z

Repeat for node_memory_utilization. Both Container Insights metrics are also published per node with the NodeName and InstanceId dimensions.

Four facts needed before one node is removed

  1. The node group’s status is ACTIVE; groups mid-update are skipped.
  2. Its desired size is above its minimum size, so there is a node the group is allowed to lose.
  3. Over 14 days, average node CPU is 40% or less and average node memory is 50% or less. Both series are required: a memory-bound group would push the survivors into out-of-memory kills on a drain.
  4. The capacity type is on-demand and an On-Demand rate is known for the worker instance type.

Node groups this rule does not touch

Spot node groups are skipped, because ZopNight holds only On-Demand rates for worker types and a removed Spot node priced at the On-Demand rate would overstate the saving. Groups already at their minimum, and groups without both metrics, produce nothing. A node group doing no work at all is covered by Running EKS Node Group Idle.

One node’s monthly cost

Terminal window
saving = On-Demand hourly rate of the worker type x 730 x 1 node
new desired size = current desired size - 1

The step is always one node. That under-claims on a very oversized group, but each later scan can take another step once the metrics show there is still room.

Scaling the node group down by one

ZopNight can apply this by setting the new desired size on the node group. By hand:

  1. Check that pods have room elsewhere, including DaemonSets and anything with node affinity.
  2. Lower the desired size: aws eks update-nodegroup-config --cluster-name prod --nodegroup-name general --scaling-config desiredSize=4
  3. Watch for pods stuck in Pending and for evictions finishing.
  4. If an autoscaler controls this group’s size, make the change in its settings too, so the two do not disagree.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·