Requests & limits
Being a good neighbor.
Lots of apps share the same machines. Each one has to leave capacity for the others. Requests and limits are the two numbers that enforce that.
First, a few plain words. A container is your app bundled with everything it needs to run, so it behaves the same wherever it lands. A cluster is a pool of machines, called nodes, that run those containers. Each container, or a small group of them, runs inside a Pod, the smallest thing Kubernetes runs and manages. Because many Pods share the same nodes, one greedy app can grab all the CPU or memory and starve its neighbors. Requests and limits stop that.
The two numbers you set
For each container you set two numbers: a request and a limit. They sound alike, but they do opposite jobs. A request is a promise the cluster makes to you. A limit is a promise you make to the cluster.
A request is what your container needs, and the scheduler treats it as reserved capacity. It sums requests on a node and only places your Pod where the remainder fits.
A limit is the most the container is allowed to use. Cross it and Kubernetes steps in; what happens depends on which resource you went over. The file below writes both numbers. It is YAML. You apply it; the cluster creates the objects.
apiVersion: apps/v1kind: Deploymentmetadata:name: webspec:replicas: 2selector:matchLabels: { app: web }template:metadata:labels: { app: web }spec:containers:- name: webimage: my-app:1.4resources:requests:cpu: 250m # a quarter of one CPU core, reserved for this containermemory: 256Mi # 256 mebibytes, reservedlimits:cpu: 500m # push past this and the container is slowed downmemory: 512Mi # push past this and the container is stopped and restarted
$ kubectl apply -f deploy.yamldeployment.apps/web created
CPU and memory are treated differently, and the reason is physical. CPU can be sliced moment by moment, so a container that passes its CPU limit is simply slowed down. That's called throttling: the app doesn't die, it runs at the speed of the cap. Memory is different. Once bytes are handed out they're gone until the app frees them, so a container that passes its memory limit gets stopped and started fresh.
$ kubectl top podsNAME CPU(cores) MEMORY(bytes)web-7d9f8c6b5-4xk2p 12m 140Miweb-7d9f8c6b5-q8m7n 9m 131Mi
That's the healthy picture, usage sitting under the limits. Push a container past its memory limit, though, and the first sign is usually a Pod that won't stay up.
$ kubectl get podsNAME READY STATUS RESTARTS AGEweb-7d9f8c6b5-4xk2p 0/1 CrashLoopBackOff 4 (18s ago) 2mweb-7d9f8c6b5-q8m7n 1/1 Running 0 2m
CrashLoopBackOff means the container keeps dying and Kubernetes keeps restarting it, waiting longer between tries. The status tells you something's wrong, not what. kubectl describe answers that; read the Last State near the top.
$ kubectl describe pod web-7d9f8c6b5-4xk2p...Last State: TerminatedReason: OOMKilledExit Code: 137Restart Count: 4
OOMKilled is short for out-of-memory killed: the container used more memory than its limit allowed, so Kubernetes stopped it. Exit Code 137 is the same story in numbers, what a process gets when it's force-stopped. That pair together is the tell. The fix is usually a higher memory limit or a patched leak, not a deleted limit.
Three tiers decide who dies first
OOMKilled raises a bigger question: when a whole node runs out of memory, which Pods die first? Kubernetes answers with a ranking. Every Pod falls into one of three quality-of-service classes, written QoS, and you never set it yourself. Kubernetes derives it from your requests and limits.
The rule is short. Set requests equal to limits for both CPU and memory on every container and the Pod is Guaranteed, the most protected, killed last. Set some requests or limits but not matching pairs and it's Burstable, the middle tier. Set nothing at all and it's BestEffort, first out the door when the node is squeezed. The more precisely you declare what you need, the safer your Pod sits.
$ kubectl get pod web-7d9f8c6b5-q8m7n -o jsonpath='{.status.qosClass}'Burstable
Our web Pod comes back Burstable, and that fits: deploy.yaml set requests but higher limits, so the two don't match. Make them identical and it prints Guaranteed; leave the resources out and it prints BestEffort, the tier you least want a real app in.
Backfill defaults with a LimitRange
On a shared cluster you can't trust everyone to remember these numbers, and a container that forgets them lands in BestEffort. A LimitRange fixes that for a whole namespace: a rule that stamps default requests and limits onto any container that arrives without its own.
apiVersion: v1kind: LimitRangemetadata:name: default-resourcesnamespace: defaultspec:limits:- type: ContainerdefaultRequest: # request stamped on if a container sets nonecpu: 250mmemory: 256Midefault: # limit stamped on if a container sets nonecpu: 500mmemory: 512Mi
$ kubectl apply -f limits.yamllimitrange/default-resources created
Now test it. Start a bare Pod with no resources block, then ask for its QoS class.
$ kubectl run probe --image=nginx --restart=Neverpod/probe created$ kubectl get pod probe -o jsonpath='{.status.qosClass}'Burstable
Without the LimitRange that Pod would've been BestEffort. The rule caught it on the way in, stamped on the default request and limit, and lifted it to Burstable, off the front of the kill list. Sane defaults for everyone, no nagging required.
Why this is worth the effort
Set both numbers early, because three things depend on them. Scheduling: with requests, the scheduler only places Pods on nodes that have room, so you avoid Pending Pods and overpacked nodes. Autoscaling: the Horizontal Pod Autoscaler (HPA) adds copies when load climbs and measures usage as a percentage of the request, so with no request it has no baseline and does nothing. Stability: a memory limit stops one container from taking the node down.
There's a balance to strike. Set requests too low and the scheduler overpacks nodes, so apps fight over scraps; set them too high and you pay for room you never touch. A safe rule: always set requests, always set a memory limit, and go easy on tight CPU limits, since throttling a healthy app just to hit a number causes more grief than it prevents.
Requests are what the scheduler uses to place pods; limits are the ceiling the runtime enforces. A pod without requests can pack densely until noisy neighbors appear. Prefer honest requests based on observed usage rather than cargo-cult numbers.
CPU limits throttle; memory limits OOM-kill. Those feel different in production. Watch for CrashLoop caused by memory limits that are too tight versus CPU throttling that just makes the app slow.
Suppose one team omits limits on a shared node pool: their leak becomes everyone's incident. Quotas and LimitRanges exist to make good-neighbor defaults enforceable.
Try this
Run a pod with explicit CPU/memory requests and limits, then inspect what the PodSpec stored and how the node allocated it.
$ kubectl apply -f - <<'EOF'apiVersion: v1kind: Podmetadata:name: sizedspec:containers:- name: cimage: nginx:1.27resources:requests:cpu: "100m"memory: "64Mi"limits:cpu: "200m"memory: "128Mi"EOFpod/sized created$ kubectl get pod sized -o jsonpath='{.spec.containers[0].resources}{"\n"}'{"limits":{"cpu":"200m","memory":"128Mi"},"requests":{"cpu":"100m","memory":"64Mi"}}$ kubectl describe pod sized | Select-String -Pattern 'Limits|Requests|QoS' -Context 0,3QoS Class: BurstableLimits:cpu: 200mmemory: 128MiRequests:cpu: 100mmemory: 64Mi$ kubectl delete pod sizedpod "sized" deleted
Takeaway
Requests place pods; limits cap them. Set both from real usage, remember memory over-limit kills while CPU throttles, and enforce defaults with LimitRange/Quota on shared clusters.
Want to feel this rather than just read about it? Deploy the file above to a small local cluster with kind or minikube (both run a tiny cluster on your laptop), run kubectl top pods, and watch usage sit under your request. Then cut the memory limit to something tiny like 20Mi, apply again, and watch the Pod flip to OOMKilled in seconds.