Kubernetes autoscaling with the Horizontal Pod Autoscaler
Scale on CPU, memory, or custom Prometheus metrics, and tune stabilization windows so the HorizontalPodAutoscaler won't flap under spiky, bursty load.
The HorizontalPodAutoscaler manifest is a dozen lines, and the two reasons it does nothing are not in it. It needs the metrics API, which means metrics-server (or a replacement) installed and answering, and it needs a CPU request on every container of the target, because utilisation is computed as usage divided by request. A Deployment without requests reports no utilisation, and the HPA sits at minReplicas while the service falls over.
kubectl top pods -n prod -l app=apiNAME CPU(cores) MEMORY(bytes)api-7d4k2 245m 128Mimetrics-server answers: prerequisite onekubectl get deploy api -n prod -o jsonpath="{range .spec.template.spec.containers[*]}{.name}: {.resources.requests.cpu}{'\n'}{end}"api: 250mlog-shipper: the sidecar has no request, so the Pod's utilisation is undefined until it gets oneThe controller reads per-Pod usage from the metrics API, divides by the requests, compares the average with the target, and scales the Deployment. Every box in the middle is a place a missing prerequisite makes the loop silently stop.
The arithmetic, and the tolerance that stops it twitching
Desired replicas are ceil(current × (current utilisation / target)), computed across all ready Pods. Three Pods at 90 % against a 70 % target want ceil(3 × 1.29) = 4. The controller ignores ratios within a tolerance of the target (10 % by default) so a 72 % reading does not add a replica, and since v1.37 that tolerance can be set per HPA in behavior.scaleUp.tolerance and behavior.scaleDown.tolerance instead of only cluster-wide. Pods that are not ready, and Pods whose metrics are missing, are handled conservatively: a missing metric counts as 100 % of target for scale-down decisions and 0 % for scale-up, so a metrics gap does not itself cause a scaling event.
apiVersion: autoscaling/v2kind: HorizontalPodAutoscalermetadata:name: apinamespace: prodspec:scaleTargetRef:apiVersion: apps/v1kind: Deploymentname: apiminReplicas: 3 # survives one node loss without dropping below 2maxReplicas: 20 # bounded by downstream connection limits, not by budget alonemetrics:- type: Resourceresource:name: cputarget:type: UtilizationaverageUtilization: 70
What the defaults already do, and what to change
Scale-up has no stabilisation window by default: when the metric says grow, the controller grows immediately, by up to 100 % of the current count or 4 Pods per 15 seconds, whichever is larger. Scale-down is the opposite: the controller looks back over a 300-second window (the --horizontal-pod-autoscaler-downscale-stabilization default) and uses the highest recommendation from that window, so replicas are removed only after five minutes of consistently lower load. A service that yo-yos is usually one where the window is too short for its traffic pattern, or where scale-up is so aggressive that the new replicas overshoot and immediately look idle.
behavior:scaleDown:stabilizationWindowSeconds: 600 # look back 10 min before removing capacitypolicies:- type: Percentvalue: 25 # remove at most a quarter of the replicas per minuteperiodSeconds: 60tolerance: 0.05 # v1.37+: react to a 5 % drift, not the cluster defaultscaleUp:stabilizationWindowSeconds: 0 # keep the default: grow immediatelypolicies:- type: Percentvalue: 100periodSeconds: 60- type: Podsvalue: 4periodSeconds: 60selectPolicy: Max
kubectl get hpa api -n prod -wNAME REFERENCE TARGETS MINPODS MAXPODS REPLICASapi Deployment/api cpu: 42%/70% 3 20 3api Deployment/api cpu: 88%/70% 3 20 4api Deployment/api cpu: 71%/70% 3 20 4kubectl describe hpa api -n prod | grep -A4 ConditionsAbleToScale True ReadyForNewScaleScalingActive True ValidMetricFoundScalingLimited False DesiredWithinRangeScalingActive False with FailedGetResourceMetric is the prerequisites failing, not the loadWhen CPU is not what users feel
CPU utilisation is a good proxy for a stateless request handler and a poor one for a queue consumer, an I/O-bound service or anything that blocks on a downstream. For those the HPA scales on a custom metric served through an adapter (the Prometheus adapter is the common one): requests per second per Pod with type: Pods, or queue depth with type: External. The metric should be one whose growth means users are waiting, and the target should be a value at which one Pod is comfortably busy, which is a number that comes from a load test rather than from a default.
metrics:- type: Podspods:metric:name: http_requests_per_second # served by the Prometheus adaptertarget:type: AverageValueaverageValue: "500"- type: Resource # keep CPU as a second signal; the HPA takes the larger resultresource:name: cputarget:type: UtilizationaverageUtilization: 70
The HPA changes the replica count; it assumes the per-Pod requests and limits are right, and it assumes new Pods can be scheduled, which is the Cluster Autoscaler's job when they cannot. The last row of the table is the reminder that scaling the tier that is not the bottleneck only moves the queue somewhere else.