# EKS Underutilized Node Group

> Finds EKS node groups with spare CPU and memory above their minimum size and prices removing one on-demand node.

Source: https://zop.dev/integrations/aws/recommendations/eks-underutilized-node-group

---

## Each extra worker node is a full EC2 bill

A managed node group is an EC2 Auto Scaling group that
[Amazon EKS manages for you](https://docs.aws.amazon.com/eks/latest/userguide/managed-node-groups.html),
and every node in it is an EC2 instance billed at its normal rate. The group's desired size is the
number of those instances running right now. If the scheduler has room to spare on every node, the
last node added is paying for nothing but headroom.

Kubernetes requests decide where pods land, but the bill follows the node count. A cluster whose
nodes average a third of their CPU and less than half their memory can usually lose a node without
any pod going unscheduled.

## Checking node-group size and node load

```bash
aws eks describe-nodegroup --cluster-name prod --nodegroup-name general \
  --query 'nodegroup.[status,capacityType,instanceTypes,scalingConfig]'

aws cloudwatch get-metric-statistics --namespace ContainerInsights \
  --metric-name node_cpu_utilization --dimensions Name=ClusterName,Value=prod \
  --statistics Average --period 86400 \
  --start-time 2026-09-11T00:00:00Z --end-time 2026-09-25T00:00:00Z
```

Repeat for `node_memory_utilization`. Both
[Container Insights metrics](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Container-Insights-metrics-EKS.html)
are also published per node with the `NodeName` and `InstanceId` dimensions.

## Four facts needed before one node is removed

1. The node group's status is `ACTIVE`; groups mid-update are skipped.
2. Its desired size is above its minimum size, so there is a node the group is allowed to lose.
3. Over 14 days, average node CPU is 40% or less and average node memory is 50% or less. Both series
   are required: a memory-bound group would push the survivors into out-of-memory kills on a drain.
4. The capacity type is on-demand and an On-Demand rate is known for the worker instance type.

## Node groups this rule does not touch

Spot node groups are skipped, because ZopNight holds only On-Demand rates for worker types and a
removed Spot node priced at the On-Demand rate would overstate the saving. Groups already at their minimum, and groups without both
metrics, produce nothing. A node group doing no work at all is covered by
<a href="https://zop.dev/integrations/aws/recommendations/running-eks-node-group-idle">Running EKS Node Group Idle</a>.

## One node's monthly cost

```text
saving = On-Demand hourly rate of the worker type x 730 x 1 node
new desired size = current desired size - 1
```

The step is always one node. That under-claims on a very oversized group, but each later scan can
take another step once the metrics show there is still room.

## Scaling the node group down by one

ZopNight can apply this by setting the new desired size on the node group. By hand:

1. Check that pods have room elsewhere, including DaemonSets and anything with node affinity.
2. Lower the desired size:
   `aws eks update-nodegroup-config --cluster-name prod --nodegroup-name general --scaling-config desiredSize=4`
3. Watch for pods stuck in `Pending` and for evictions finishing.
4. If an autoscaler controls this group's size, make the change in its settings too, so the two do
   not disagree.

**Warning**
AWS states that pod disruption budgets are not respected when the desired node count is
reduced, and a node is terminated after 15 minutes even if pods have not all been evicted. Make sure
critical workloads have spare replicas first.
