Skip to main content
idle · aws

Degraded or zero-size EKS node groups whose clusters still pay for orphaned volumes and load balancers

resource types
1
rule IDs covered
1
severity
medium

What does ZopNight detect here?

ZopNight flags EKS managed node groups in `DEGRADED` status or scaled to a desired size of 0 when their cluster still holds orphaned attachments: Released or Available PersistentVolumes backed by EBS volumes, and load balancers with no healthy targets. The saving is the priced monthly cost of those attachments, counted once per cluster.

Signal and threshold

How ZopNight evaluates Degraded or zero-size EKS node groups whose clusters still pay for orphaned volumes and load balancers.
Field Value
Rule IDsRC-043
Categoryidle
Severitymedium
Metricnone — pure configuration read
ThresholdDEGRADED or desiredSize 0, orphaned attachments
SourceZopNight
Permissions usedeks:ListNodegroups · eks:DescribeNodegroup · ec2:DescribeVolumes · elasticloadbalancing:DescribeLoadBalancers · elasticloadbalancing:DescribeTargetHealth

Scaling nodes away does not delete what the pods left behind

When the nodes of an EKS node group go away, two kinds of AWS resources created on behalf of the cluster can keep billing. The first is storage. The Kubernetes volume docs explain that under the Retain reclaim policy, deleting a PersistentVolumeClaim leaves the PersistentVolume in place as “released”, with the underlying EBS volume still provisioned. The second is networking: a Service of type LoadBalancer, as the EKS load balancing guide describes, creates an AWS Classic or Network Load Balancer, which keeps its hourly charge after the backends disappear.

A node group that is broken or scaled to zero is the clearest sign nobody is watching those leftovers.

Spotting an abandoned node group and its leftovers

Terminal window
aws eks describe-nodegroup --cluster-name old-cluster --nodegroup-name workers \
--query 'nodegroup.[status,scalingConfig.desiredSize,health.issues]'
kubectl get pv -o wide | grep -E 'Released|Available'
kubectl get svc -A --field-selector spec.type=LoadBalancer

Node group status values include ACTIVE and DEGRADED. For each load balancer hostname, check aws elbv2 describe-target-health for healthy targets (or aws elb describe-instance-health for a Classic Load Balancer).

Two signals plus a cost

The node group qualifies when its status is DEGRADED, or when its desired size is 0. Then ZopNight looks across the cluster, not just the node group, and joins what Kubernetes reports to what AWS bills:

  • a PersistentVolume in Released or Available state whose CSI volume handle matches an EBS volume ZopNight has priced;
  • a LoadBalancer Service whose ingress hostname matches an AWS load balancer that has no healthy target.

At least one priced attachment must be found, or the node group raises nothing.

How double counting is avoided

The node group itself has no cost of its own; its EC2 nodes are billed and priced separately. The orphaned attachments are summed per cluster and assigned to one node group, so a cluster with three abandoned node groups gets one finding, not three. A volume or load balancer referenced by two abandoned clusters is counted once. Attachments counted here are skipped by Orphaned EBS Volume and Idle Load Balancer, so the same dollars never appear twice. An active, healthy node group with a non-zero desired size is never flagged, whatever else is orphaned in the cluster.

The saving is the leftover attachments

Terminal window
saving = monthly cost of released or available EBS-backed volumes
+ monthly cost of load balancers with no healthy targets

Cleaning up the cluster

  1. Confirm no workload is expected to come back on this node group.
  2. Snapshot any released volume whose data might be needed, then delete the PV and the EBS volume.
  3. Delete LoadBalancer Services that no longer front anything, which removes their load balancers.
  4. Delete the empty node group: aws eks delete-nodegroup --cluster-name old-cluster --nodegroup-name workers

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·