Skip to main content
idle · aws

Active EKS managed node groups whose nodes ran under 3% CPU for 30 days

resource types
1
rule IDs covered
1
severity
high

What does ZopNight detect here?

ZopNight flags an active, non-Spot EKS managed node group with at least one desired node when Container Insights `node_cpu_utilization` stays under 3% on average and at peak for 30 days and average node memory is at most 50%. Worker nodes bill as EC2 instances, so the saving is the node price times 730 hours times the desired count.

Signal and threshold

How ZopNight evaluates Active EKS managed node groups whose nodes ran under 3% CPU for 30 days.
Field Value
Rule IDsRC-043b
Categoryidle
Severityhigh
Metricnode_cpu_utilization, node_memory_utilization
ThresholdCPU under 3% (average and peak); memory average at most 50%
Evaluation window30d
SourceZopNight
Permissions usedeks:ListNodegroups · eks:DescribeNodegroup · cloudwatch:GetMetricStatistics

Nodes cost the same empty as full

Amazon EKS pricing is explicit that on top of the cluster fee you pay for the resources your worker nodes use, such as EC2 instances, EBS volumes and public IPv4 addresses. A managed node group with a desired size of three is three EC2 instances billed by the hour, whatever the pods on them are doing.

Node groups outlive the workloads they were created for: a team moves to a new group with a different instance type, or the application moves to another cluster, and the old group keeps its desired count.

Looking at node utilization

Container Insights publishes node_cpu_utilization and node_memory_utilization in the ContainerInsights namespace; the EKS metrics table lists ClusterName, NodeName and InstanceId as dimensions. ZopNight instead queries each node group on the NodegroupName and ClusterName dimensions, a pair that table does not list:

Terminal window
aws eks list-nodegroups --cluster-name my-cluster
aws cloudwatch get-metric-statistics \
--namespace ContainerInsights --metric-name node_cpu_utilization \
--dimensions Name=ClusterName,Value=my-cluster Name=NodegroupName,Value=my-nodes \
--start-time 2026-08-26T00:00:00Z --end-time 2026-09-25T00:00:00Z \
--period 86400 --statistics Average Maximum

These per-node-group metrics need the Amazon CloudWatch Observability add-on (Container Insights with enhanced observability) on the cluster. AWS’s metric tables do not list NodegroupName as a dimension, so where your setup does not publish it, ZopNight’s query returns nothing and no finding is raised.

CPU says idle and memory agrees

The node group must be active with a desired size of at least 1, and its capacity type must not be Spot. Its node CPU series needs 30 days of coverage with both the average and the peak under 3%. CPU alone could miss a group holding caches or JVM heaps in memory, so the node memory series must also be present with 30 days of coverage and an average of 50% or less. A memory-heavy group is left for a trimming decision instead of a full teardown.

Groups that are skipped

Without Container Insights data there is no evidence, so no finding. Spot node groups are skipped because ZopNight can only price them at the On-Demand rate, which would overstate the saving. A group with no desired nodes is covered by EKS Abandoned Node Group, and a group that is busy but oversized by EKS Underutilized Node Group. An instance type with no rate produces nothing.

Pricing the node hours

Terminal window
saving = On-Demand hourly rate for the instance type x 730 x desired nodes
cost after fix = 0

Draining and removing the group

  1. Check for batch, seasonal or node-selector workloads that target this group.
  2. Review PodDisruptionBudgets so a drain does not block.
  3. Pause it with aws eks update-nodegroup-config --cluster-name my-cluster --nodegroup-name my-nodes --scaling-config minSize=0,maxSize=1,desiredSize=0, or remove it with aws eks delete-nodegroup.
  4. Clean up any load balancers or EBS volumes the group’s workloads left behind.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·