Kubernetes autoscaling with the Horizontal Pod Autoscaler

Scale on CPU, memory, or custom Prometheus metrics, and tune stabilization windows so the HorizontalPodAutoscaler won't flap under spiky, bursty load.

Dec 3, 2024·Updated ·4 min readIntermediate·By SecOpsLog · documentation-verified

The HorizontalPodAutoscaler manifest is a dozen lines, and the two reasons it does nothing are not in it. It needs the metrics API, which means metrics-server (or a replacement) installed and answering, and it needs a CPU request on every container of the target, because utilisation is computed as usage divided by request. A Deployment without requests reports no utilisation, and the HPA sits at minReplicas while the service falls over.

bash — the two prerequisites, checked in order
kubectl top pods -n prod -l app=api
NAME CPU(cores) MEMORY(bytes)
api-7d4k2 245m 128Mi
metrics-server answers: prerequisite one
kubectl get deploy api -n prod -o jsonpath="{range .spec.template.spec.containers[*]}{.name}: {.resources.requests.cpu}{'\n'}{end}"
api: 250m
log-shipper:
the sidecar has no request, so the Pod's utilisation is undefined until it gets one
Metrics to replicas

The controller reads per-Pod usage from the metrics API, divides by the requests, compares the average with the target, and scales the Deployment. Every box in the middle is a place a missing prerequisite makes the loop silently stop.

Metrics Server CPU / memory HPA controller targetUtilization × requests Deployment replicas N → M ! No requests = stuck utilization = usage / request missing request → 0% HPA sits at minReplicas 1 Scale the Deployment never scale Pods directly min / max + behavior windows stabilization avoids flapping 2 Verify kubectl top + HPA TARGETS load test until scale-up alert when pinned at max HPA reads metrics and patches Deployment.spec.replicas. Requests make utilization real. metrics → HPA → Deployment replicas → Pods
Metrics — serverHPA — decideScale — Deployment

The arithmetic, and the tolerance that stops it twitching

Desired replicas are ceil(current × (current utilisation / target)), computed across all ready Pods. Three Pods at 90 % against a 70 % target want ceil(3 × 1.29) = 4. The controller ignores ratios within a tolerance of the target (10 % by default) so a 72 % reading does not add a replica, and since v1.37 that tolerance can be set per HPA in behavior.scaleUp.tolerance and behavior.scaleDown.tolerance instead of only cluster-wide. Pods that are not ready, and Pods whose metrics are missing, are handled conservatively: a missing metric counts as 100 % of target for scale-down decisions and 0 % for scale-up, so a metrics gap does not itself cause a scaling event.

hpa.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api
namespace: prod
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api
minReplicas: 3 # survives one node loss without dropping below 2
maxReplicas: 20 # bounded by downstream connection limits, not by budget alone
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70

What the defaults already do, and what to change

Scale-up has no stabilisation window by default: when the metric says grow, the controller grows immediately, by up to 100 % of the current count or 4 Pods per 15 seconds, whichever is larger. Scale-down is the opposite: the controller looks back over a 300-second window (the --horizontal-pod-autoscaler-downscale-stabilization default) and uses the highest recommendation from that window, so replicas are removed only after five minutes of consistently lower load. A service that yo-yos is usually one where the window is too short for its traffic pattern, or where scale-up is so aggressive that the new replicas overshoot and immediately look idle.

hpa-behavior.yaml
behavior:
scaleDown:
stabilizationWindowSeconds: 600 # look back 10 min before removing capacity
policies:
- type: Percent
value: 25 # remove at most a quarter of the replicas per minute
periodSeconds: 60
tolerance: 0.05 # v1.37+: react to a 5 % drift, not the cluster default
scaleUp:
stabilizationWindowSeconds: 0 # keep the default: grow immediately
policies:
- type: Percent
value: 100
periodSeconds: 60
- type: Pods
value: 4
periodSeconds: 60
selectPolicy: Max
bash — watching a scale-up, and reading the conditions
kubectl get hpa api -n prod -w
NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS
api Deployment/api cpu: 42%/70% 3 20 3
api Deployment/api cpu: 88%/70% 3 20 4
api Deployment/api cpu: 71%/70% 3 20 4
kubectl describe hpa api -n prod | grep -A4 Conditions
AbleToScale True ReadyForNewScale
ScalingActive True ValidMetricFound
ScalingLimited False DesiredWithinRange
ScalingActive False with FailedGetResourceMetric is the prerequisites failing, not the load

When CPU is not what users feel

CPU utilisation is a good proxy for a stateless request handler and a poor one for a queue consumer, an I/O-bound service or anything that blocks on a downstream. For those the HPA scales on a custom metric served through an adapter (the Prometheus adapter is the common one): requests per second per Pod with type: Pods, or queue depth with type: External. The metric should be one whose growth means users are waiting, and the target should be a value at which one Pod is comfortably busy, which is a number that comes from a load test rather than from a default.

hpa-custom.yaml (metrics section)
metrics:
- type: Pods
pods:
metric:
name: http_requests_per_second # served by the Prometheus adapter
target:
type: AverageValue
averageValue: "500"
- type: Resource # keep CPU as a second signal; the HPA takes the larger result
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
Alert on ScalingLimited, and size maxReplicas from the database
An HPA pinned at maxReplicas with the target still exceeded means the service is under-provisioned by the one number a human set. Alert on ScalingLimited being true, and size maxReplicas from what the database and downstream connection pools can absorb, because those are the limits the autoscaler will otherwise find for you.
Why an HPA does nothing
Symptom
TARGETS shows <unknown>/70%
REPLICAS stuck at minReplicas under load
Replicas oscillate every few minutes
Scaled to max, still slow
Cause
metrics-server missing or a container without a CPU request
utilisation undefined (see above) or target too high
scale-down window too short for the traffic pattern
the bottleneck is not in the Pods: database, downstream, node capacity

The HPA changes the replica count; it assumes the per-Pod requests and limits are right, and it assumes new Pods can be scheduled, which is the Cluster Autoscaler's job when they cannot. The last row of the table is the reminder that scaling the tier that is not the bottleneck only moves the queue somewhere else.

Related posts

Quick reference