Kubernetes requests and limits: stop OOMKills, noisy pods
Set CPU and memory requests and limits that keep the scheduler honest, stop one pod from starving the node, and avoid needless OOMKills under load.
kubectl get pod api-7c9 -n shopapi-7c9 0/1 OOMKilled 6 14mkubectl top pod -n shop -l app=api --containersPOD NAME CPU(cores) MEMORY(bytes)api-7c9 api 140m 121Misteady state is already at 95 % of the limit; the first request burst is the kill, not a bugkubectl get pod leaky -n shopleaky 0/1 OOMKilled 1 (3s ago) 4skubectl get pod leaky -n shop -o jsonpath="{.status.containerStatuses[0].lastState.terminated.reason} {.spec.containers[0].resources.limits.memory}"OOMKilled 64Mithe kernel kills the container the moment it crosses the limit; the default restartPolicy brings it back to be killed again. The application log ends mid-allocation, with no error of its ownTwo fields on a container decide two different things, and most resource incidents come from treating them as one. Requests are what the scheduler reserves: they decide which node a Pod lands on and whether it fits at all. Limits are what the kubelet enforces at runtime: a container that exceeds its memory limit is killed by the kernel, and one that exceeds its CPU limit is throttled. A request set from a guess produces Pods that never schedule or nodes that are overcommitted; a limit set from a guess produces the restart loop above.
What each field does, and the QoS class you get
Requests, limits and Quality of Service
| You set | Scheduler | Kubelet at runtime | QoS class | Evicted under node memory pressure |
|---|---|---|---|---|
| nothing | places the Pod anywhere | no ceiling; competes for what is left | BestEffort | first |
| requests only, or limits above requests | reserves the request | memory above the limit is OOM-killed; CPU above the limit is throttled | Burstable | after BestEffort, ordered by how far usage exceeds the request |
| requests equal to limits for every container, CPU and memory | reserves exactly the limit | the same ceilings | Guaranteed | last |
The class is derived, not declared, and the derivation is strict: one container with a CPU limit above its request drops the whole Pod to Burstable. The reason to want Guaranteed for a database or a queue consumer is the last column; the reason not to force it on everything is that requests equal to limits leave no room for the burst that Burstable exists to absorb.
kubectl get pod besteffort burstable guaranteed -n shop -o jsonpath="{range .items[*]}{.metadata.name}: {.status.qosClass}{'\n'}{end}"besteffort: BestEffortburstable: Burstableguaranteed: Guaranteedno resources at all is BestEffort; requests with a higher memory limit is Burstable; requests equal to limits for cpu and memory is GuaranteedA CPU limit throttles; a memory limit kills
The two limits fail differently. Memory is not compressible, so the only enforcement is termination: the container is OOM-killed, lastState says OOMKilled, and the application log usually ends mid-sentence. CPU is compressible, so the limit is enforced by CFS quota: the container keeps running but is paused whenever it has used its share of each scheduling period, which shows up as latency with no error anywhere. That asymmetry is why many teams set memory limits on everything and CPU limits on almost nothing: a memory limit protects the node from a leak, while a CPU limit only ever slows the workload down and never protects a neighbour that the request did not already protect. The container_cpu_cfs_throttled_periods_total metric is how you find out a CPU limit is the cause of a latency complaint.
resources:requests:cpu: 100m # near steady-state usage: what the scheduler reservesmemory: 256Milimits:memory: 256Mi # equal to the request: memory is never overcommitted for this container# no cpu limit: bursts use idle node capacity instead of being throttled
Numbers come from measurements, and they can change without a restart
Deploy with requests only, run through a normal day and a peak, then read kubectl top or the metrics stack for the p50 and the p99 of usage. Requests go near the p50 so the scheduler packs efficiently; the memory limit goes above the observed peak with headroom for garbage-collection spikes. On Kubernetes v1.35 and later the correction does not require a rollout: the resize subresource changes a running container's requests and limits in place, subject to the container's resizePolicy (memory decreases may require a restart, which the policy declares).
kubectl get pod api -n shop -o jsonpath="{.status.containerStatuses[0].resources.limits.memory} restarts={.status.containerStatuses[0].restartCount}"128Mi restarts=0kubectl patch pod api -n shop --subresource resize --patch '{"spec":{"containers":[{"name":"api","resources":{"requests":{"memory":"256Mi"},"limits":{"memory":"256Mi"}}}]}}'pod/api patchedkubectl get pod api -n shop -o jsonpath="{.status.containerStatuses[0].resources.limits.memory} restarts={.status.containerStatuses[0].restartCount}"256Mi restarts=0the running container has the new limit and was not restarted (restartCount unchanged); the Deployment template still needs the same change so the next rollout keeps itThe Vertical Pod Autoscaler in Off (recommendation) mode does the measuring continuously and writes its recommendation into the VPA object's status without touching Pods, which is a safe way to keep numbers current across many Deployments and to notice when a workload's profile drifts. Its Auto mode applies the recommendation itself; with in-place resize available, that is less disruptive than it was, but it is still a second controller changing your spec, and a change nobody reviewed.
Requests and limits size one container. What keeps a team from taking the whole node is a ResourceQuota on the namespace and a LimitRange that supplies defaults for Pods that declare nothing, and what keeps the scheduler from silently failing is reading the FailedScheduling event rather than the Pod spec. When the right answer is more replicas rather than bigger ones, the HPA needs exactly these CPU requests to compute utilisation.