# HPA Cannot Scale

> Flags HPAs whose maxReplicas is 1, which can never add a replica and only give the impression of autoscaling.

Source: https://zop.dev/integrations/kubernetes/recommendations/hpa-cannot-scale

---

## An autoscaler with no room to scale

The
[HPA API reference](https://kubernetes.io/docs/reference/kubernetes-api/workload-resources/horizontal-pod-autoscaler-v2/)
defines `maxReplicas` as the upper limit the autoscaler can scale up to, and says it cannot be less
than `minReplicas`, which defaults to 1. Set the maximum to 1 and the range collapses to a single
value. The controller still runs, still reads metrics and still reports conditions, but every
calculation ends at one pod.

One common route is a template: a chart ships with `maxReplicas: 1` as a safe default for dev,
and the value is never overridden for production. The service then has one pod, no redundancy and
no ability to absorb a spike, while dashboards and reviews see an HPA and assume the scaling
question is handled.

## Listing HPAs capped at one

```bash
kubectl get hpa -A -o json | jq -r '
  .items[]
  | select(.spec.maxReplicas == 1)
  | "\(.metadata.namespace)/\(.metadata.name) min=\(.spec.minReplicas // 1) target=\(.spec.scaleTargetRef.kind)/\(.spec.scaleTargetRef.name)"'
```

## Exactly one, nothing more

ZopNight reads `maxReplicas` from the HPA and fires when it equals 1. There is no metric,
threshold or window beyond that single value, and the finding clears when the maximum is raised.
If the value was not reported, the HPA is skipped.

## Cases this check leaves alone

Pinned HPAs with minimum equal to a larger maximum, such as 3 and 3, are the subject of
[HPA pinned](https://zop.dev/integrations/kubernetes/recommendations/hpa-pinned), which deliberately skips the
case of 1 so the two checks do not overlap. HPAs capped higher that run out of headroom appear
under [HPA at max capacity](https://zop.dev/integrations/kubernetes/recommendations/hpa-at-max-capacity). The
check does not look at what the HPA targets, so an HPA pointing at a Deployment that no longer
exists is flagged the same way.

## A reliability finding without a price

This medium-severity reliability recommendation carries no savings estimate. Raising the maximum
does not cost anything until load actually calls for more pods.

## Opening up the range

1. Decide the most replicas the service and its dependencies can handle at peak.
2. Raise the maximum in the chart or manifest that owns the HPA, or patch it directly:
   `kubectl patch hpa <name> -p '{"spec":{"maxReplicas":5}}'`.
3. Consider raising `minReplicas` to 2 for production services so a single pod failure does not
   take the service down.
4. Confirm the target workload has CPU or memory requests, since utilization targets are
   calculated against requests.
5. If a fixed single pod is really intended, delete the HPA so the manifest says what it does.
