Skip to main content
reliability · kubernetes

HorizontalPodAutoscalers with an empty metrics list

resource types
1
rule IDs covered
3
severity
medium

What does ZopNight detect here?

ZopNight flags a Kubernetes HorizontalPodAutoscaler on EKS, GKE or AKS whose collected spec lists no entries under `metrics`. An HPA scales only on the metrics it names, so an empty list leaves the scaling signal undefined; the `autoscaling/v2` API substitutes an 80% average CPU target, which may not be what the owner intended.

Signal and threshold

How ZopNight evaluates HorizontalPodAutoscalers with an empty metrics list.
Field Value
Rule IDsRC-1746 · RC-1846 · RC-1946
Categoryreliability
Severitymedium
Metricspec.metrics
Thresholdabsent or empty list
SourceZopNight
Permissions usedlist horizontalpodautoscalers.autoscaling

What an HPA with no named metric scales on

Every scaling decision is desiredReplicas = ceil(currentReplicas x currentMetricValue / desiredMetricValue), per the horizontal pod autoscaling page. Without a metric there is no ratio to compute.

Kubernetes papers over the gap in one place. The HPA v2 API reference says that if metrics is not set, the default metric is 80% average CPU utilization, and the API defaulting code fills that in when the object is stored. That default is a guess on the owner’s behalf. A memory-bound or queue-driven service scaled on 80% CPU either never scales or scales at the wrong moment, and CPU utilization can only be computed when the target’s containers set CPU requests.

Checking the stored metric list

Terminal window
kubectl get hpa -A -o json | jq -r '
.items[]
| select((.spec.metrics // []) | length == 0)
| "\(.metadata.namespace)/\(.metadata.name)"'

Also look at one HPA in full with kubectl get hpa <name> -o yaml. If the output shows a CPU target you never wrote, the API default is in effect and the owner should decide whether 80% CPU is the signal they want.

An absent or empty list is the trigger

ZopNight reads the HPA’s metric list as it collected it and fires when the list is missing or has no entries. It does not read the HPA’s conditions or current metric values, and there is no window. Once a metric is present in what ZopNight collects, the finding clears on the next evaluation.

Why the finding is rare in practice

Because the check reads the collected specification, the result depends on what the API returns. Read through autoscaling/v2, an HPA that omitted metrics comes back with the 80% CPU default filled in, so it has a metric and this rule does not flag it. The rule therefore does not catch HPAs that silently rely on that default. It also does not verify whether the target’s containers have CPU requests, so an HPA with a CPU target that cannot be computed is not caught here; the requests side is covered by Missing CPU/memory requests.

Reliability finding with no dollar figure

This medium-severity recommendation has no savings estimate. An HPA without a deliberate signal either wastes replicas at quiet times or fails to add them at busy ones, and neither shows up until traffic changes.

Naming the metric explicitly

  1. Decide what load looks like for the service: CPU, memory, requests per second or queue depth.
  2. Add the metric to the HPA spec, for example a Resource metric for cpu with target.type: Utilization and averageUtilization: 70.
  3. For custom or external metrics, confirm the metrics adapter serving custom.metrics.k8s.io or external.metrics.k8s.io is installed.
  4. Run kubectl describe hpa <name> and check ScalingActive is True with a valid metric found.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·