Skip to main content
reliability · kubernetes

HPAs held at maxReplicas while still asking for more pods

resource types
1
rule IDs covered
3
severity
high

What does ZopNight detect here?

ZopNight flags a Kubernetes HorizontalPodAutoscaler on EKS, GKE or AKS when `currentReplicas` equals `maxReplicas` and the HPA reports `ScalingLimited` as True with reason `TooManyReplicas`. The autoscaler has calculated that it needs more pods than its ceiling allows, so demand above that point goes unserved.

Signal and threshold

How ZopNight evaluates HPAs held at maxReplicas while still asking for more pods.
Field Value
Rule IDsRC-1717 · RC-1817 · RC-1917
Categoryreliability
Severityhigh
Metricstatus.currentReplicas and ScalingLimited condition
Thresholdcurrent equals max, ScalingLimited True with TooManyReplicas
SourceZopNight
Permissions usedlist horizontalpodautoscalers.autoscaling

The ceiling is now the bottleneck

The autoscaler computes a desired replica count from the ratio of current to target metric values, per the horizontal pod autoscaling algorithm, then clamps it between minReplicas and maxReplicas. When the clamp bites, the HPA sets a ScalingLimited condition. The HPA walkthrough describes that condition as meaning the desired scale was capped by the minimum or maximum, and a hint that you may want to change those bounds. In the Kubernetes controller source the reason for a cap at the top is TooManyReplicas.

At that point the pods that exist absorb all the extra load. Latency rises, CPU-bound pods queue work, and requests start failing if the load keeps growing, while the HPA can only report that it would like to help.

Finding HPAs stuck at the top

Terminal window
kubectl get hpa -A -o json | jq -r '
.items[]
| select(.status.currentReplicas == .spec.maxReplicas
and any(.status.conditions[]?; .type == "ScalingLimited"
and .status == "True" and .reason == "TooManyReplicas"))
| "\(.metadata.namespace)/\(.metadata.name) \(.status.currentReplicas)/\(.spec.maxReplicas) desired=\(.status.desiredReplicas)"'

kubectl describe hpa <name> shows the same condition with its message.

Two signals must agree

ZopNight needs the maximum, current and desired replica counts on the HPA; if any is missing, the HPA is skipped. It fires only when current replicas equal the maximum and the HPA’s own conditions include ScalingLimited, status True, reason TooManyReplicas. Sitting at the maximum is not enough by itself: an HPA can rest at its ceiling with load that exactly fits, and the condition is what shows the controller wanted more. The check reads the state at evaluation time and has no window.

What it leaves to other checks

An HPA whose maximum is 1 cannot scale at all and is reported by HPA cannot scale; one with minimum equal to maximum appears under HPA pinned. An HPA held at its minimum is not used as a rightsizing signal; the CPU over-provisioned check compares requests with measured CPU usage instead.

Risk to availability, no saving

This high-severity reliability finding has no dollar amount. Raising the ceiling usually costs more, which is the point: the workload needs the capacity.

Giving the autoscaler room

  1. Confirm the demand is real by comparing current and target metric values in kubectl describe hpa <name>.
  2. Raise the ceiling, for example kubectl patch hpa <name> -p '{"spec":{"maxReplicas":20}}'.
  3. Check the nodes can hold the extra pods, or that a cluster autoscaler can add nodes; otherwise new replicas will sit Pending.
  4. Check the namespace ResourceQuota has room for the new maximum.
  5. If per-pod efficiency is the problem, profile the app before adding more replicas.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·