ResourceQuotas running at 85% or more of a hard limit
What does ZopNight detect here?
ZopNight warns when a Kubernetes ResourceQuota on EKS, GKE or AKS has any resource between 85% and just under 100% of its hard limit. That band is where ordinary operations, such as a rolling update's 25% surge or one HPA scale-up, can push the namespace over the line and start getting 403 Forbidden.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1756 · RC-1856 · RC-1956 |
| Category | governance |
| Severity | medium |
| Metric | status.used / status.hard |
| Threshold | 85% to below 100% on any resource |
| Source | ZopNight |
| Permissions used | list resourcequotas |
Where it applies
Headroom disappears faster than usage suggests
Steady-state usage is not the peak a quota has to absorb. Two routine events briefly need more than the namespace normally holds:
- Rolling updates. A Deployment on the default
RollingUpdatestrategy may run up to 25% more pods than desired while it swaps versions, according to the Deployments documentation. Those surge pods count againstpods,requests.cpuandrequests.memorylike any other. - Autoscaling. An HPA adding replicas during a traffic spike needs quota at the worst possible moment.
When the extra pods do not fit, the API server refuses them with 403 Forbidden, as described in the resource quotas documentation. A namespace at 90% of its pod count can therefore run fine for weeks and then fail its next release.
Measuring how close each quota is
kubectl get resourcequota -A \ -o custom-columns=NAMESPACE:.metadata.namespace,NAME:.metadata.name,USED:.status.used,HARD:.status.hardkubectl describe resourcequota -n <namespace> prints the same data as a table per resource,
which is easier to read for CPU and memory values with units.
Why 85% is the trigger
ZopNight compares used with hard for every resource that the quota reports on both sides,
converting quantities such as 500m or 8Gi so the ratio is meaningful. It takes the highest
ratio in the quota and fires when that ratio is at least 0.85 and below 1.0. The finding quotes the
percentage it found, so a quota at 92% says 92%. There is no time window: the ratio is recomputed
each evaluation, and a quota that drops below 85% clears on its own.
As a worked example, a quota with pods: "20" and 18 pods running is at 90%. A four-replica
Deployment in that namespace needs one surge pod to roll, which takes it to 19; two such rollouts
at once would need 20, and a third would be refused.
When it stays quiet
At 100% or above, the quota is reported by ResourceQuota exhausted rather than here, so the two findings never overlap. Entries with a hard limit of zero, entries with no matching usage value and values that cannot be parsed are left out of the ratio.
An early warning, not a saving
The recommendation carries no savings estimate. It exists to give the namespace owner time to act before a deploy is blocked.
Restoring headroom
- Identify the resource near its cap from the
describeoutput. - Decide whether usage is legitimate growth or leftovers such as failed Jobs and idle Deployments, and clean up the leftovers first.
- If growth is real, raise
spec.hardfor that resource with enough margin for surge pods and the HPA’smaxReplicas. - Revisit the quota after the next release to confirm the rollout fit comfortably.