cgroups v2 and resource containment

Memory, CPU and PID limits, OOM kills and where to read them.

Advanced14 min · lesson 2 of 24
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (1 files, 1 KB): cgroups.tar.gz. The lab VM shares no folders with your computer, so fetch them inside the VM: cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-int/cgroups.tar.gz && tar -xzf cgroups.tar.gz, which creates ~/lab/cgroups/. SHA-256: 289e8645436ee1b50f59d8352704e08222a9b9e76d465f449e0aea547d458f75

A batch worker has --memory 64m. After a run, docker inspect reports OOMKilled=true and ExitCode=0, and the engineer on call files it as a Docker bug, because a container that was killed for memory cannot have exited cleanly. Both values are correct. Reading them correctly needs the mechanism underneath: control groups, or cgroups, the kernel feature that counts and limits what a group of processes uses. Namespaces, from the previous lesson, decide what a container can see. cgroups decide how much memory, CPU time and how many processes it can take.

Everything runs on the main lab VM and stays inside lab-* containers. One small image, lab-stress:1, provides a load generator: its Dockerfile is in the lesson files (copy it to ~/lab/cgroups).

Dockerfile
# lab-stress:1, a small load generator for the cgroups lesson (stress-ng from Alpine's community repo)
FROM alpine:3.22
RUN apk add --no-cache stress-ng
ENTRYPOINT ["stress-ng"]

stress-ng is maintained and packaged in Alpine. The progrium/stress image that older tutorials use is stored in the Docker schema 1 manifest format, which containerd 2.1 and later refuse to pull, so it no longer works on Docker 29.

One tree, one scope per container

Build the image in ~/lab/cgroups:

ubuntu@secopslog-docker:~/lab/cgroups · Docker 29.8.2
$ docker build -q -t lab-stress:1 .
sha256:41c216405d17e510a4151f06a5bf2a7f28909c2a8dcd1e0d974ef40c21cf4dc9

Then ask Docker and the kernel which cgroup setup this host has:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker info --format 'cgroup v{{.CgroupVersion}}, driver {{.CgroupDriver}}' stat -fc %T /sys/fs/cgroup free -h | grep -E '^(Mem|Swap)'
cgroup v2, driver systemd cgroup2fs Mem: 5.8Gi 826Mi 381Mi 1.4Mi 4.8Gi 5.0Gi Swap: 0B 0B 0B

This track uses cgroup v2 only: one unified hierarchy mounted at /sys/fs/cgroup (filesystem type cgroup2fs), which current Ubuntu, Debian, Fedora and RHEL releases use by default. driver systemd means Docker asks systemd to create each container's cgroup as a transient unit called docker-<container id>.scope under system.slice, so systemd's tools see containers like any other unit. The VM has no swap (Swap: 0B), which matters for the swap section below.

Limits are files the kernel reads

Start a container with a memory limit, a CPU limit and a process limit, then read its cgroup directory on the host:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-web --memory 256m --cpus 1.5 --pids-limit 200 nginx:1.30-alpine
e51aaaeecb316bbc22e1413abe2cd85248b1402888a490a807086479421b87ba
$ CG=/sys/fs/cgroup/system.slice/docker-$(docker inspect -f '{{.Id}}' lab-web).scope cd "$CG" && grep -H . memory.max memory.swap.max memory.current cpu.max pids.max pids.current
memory.max:268435456 memory.swap.max:268435456 memory.current:5406720 cpu.max:150000 100000 pids.max:200 pids.current:5
$ docker inspect -f 'Memory={{.HostConfig.Memory}} MemorySwap={{.HostConfig.MemorySwap}} NanoCpus={{.HostConfig.NanoCpus}} PidsLimit={{.HostConfig.PidsLimit}}' lab-web
Memory=268435456 MemorySwap=536870912 NanoCpus=1500000000 PidsLimit=200

memory.current is what the group uses now (about 5 MiB for an idle nginx; yours will differ) and pids.current counts its tasks, the master and four workers. The swap line needs care. cgroup v2 has separate files for memory and swap. Docker's API still takes one combined number, MemorySwap, which defaults to twice --memory when you do not set it (536870912 above), and the runtime writes the difference into memory.swap.max. So a plain --memory 256m allows up to another 256 MiB of swap on a host that has swap. The same values also appear on the systemd unit, because systemd created the scope:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ ID=$(docker inspect -f '{{.Id}}' lab-web) cat /proc/$(docker inspect -f '{{.State.Pid}}' lab-web)/cgroup systemctl show docker-$ID.scope -p Slice -p MemoryMax -p CPUQuotaPerSecUSec -p TasksMax
0::/system.slice/docker-e51aaaeecb316bbc22e1413abe2cd85248b1402888a490a807086479421b87ba.scope Slice=system.slice CPUQuotaPerSecUSec=1.500000s MemoryMax=268435456 TasksMax=200
$ systemd-cgls --no-pager -u docker-$(docker inspect -f '{{.Id}}' lab-web).scope
Unit docker-e51aaaeecb316bbc22e1413abe2cd85248b1402888a490a807086479421b87ba.scope (/system.slice/docker-e51aaaeecb316bbc22e1413abe2cd85248b1402888a490a807086479421b87ba.scope): ├─276896 nginx: master process nginx -g daemon off; ├─276970 nginx: worker process ├─276971 nginx: worker process ├─276972 nginx: worker process └─276973 nginx: worker process

/proc/<pid>/cgroup holds a single 0:: line on cgroup v2: the process's path in the unified tree. systemd-cgls -u prints the processes in one unit. The full tree (systemd-cgls without arguments) also shows containerd and the shims inside containerd.service and dockerd inside docker.service, separate from the container scopes.

What the container sees, and what it does not

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker exec lab-web sh -c 'cat /sys/fs/cgroup/memory.max; grep MemTotal /proc/meminfo; nproc'
268435456 MemTotal: 6039728 kB 4

The container reads its own limit from /sys/fs/cgroup/memory.max, because its cgroup namespace shows its own cgroup as the root. /proc/meminfo and nproc are not cgroup-aware: they report the VM's 6 GB and 4 CPUs. A program that sizes a heap, a cache or a thread pool from those numbers plans for resources it will never get and is killed or throttled when it grows into them. Current Java (container support is on by default since JDK 10) and Go (GOMAXPROCS follows the CPU limit since Go 1.25) read the cgroup files; check what your runtime and your own scripts read.

Setting --memory-swap equal to --memory gives the container no swap at all:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm --memory 256m --memory-swap 256m alpine:3.22 cat /sys/fs/cgroup/memory.swap.max
0

On a host with swap, a container that reaches memory.max without that setting starts swapping instead of being killed. memory.current never goes above memory.max; the swapped pages appear in memory.swap.current, so docker stats keeps showing usage at or under the limit while the container slows down. Because the lab VM has no swap, the swap allowance here is a number in a file with no effect, and the lesson does not demonstrate swapping.

A container with no limits has no memory or CPU cap, but it does have a process cap:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm alpine:3.22 sh -c 'cd /sys/fs/cgroup && grep -H . memory.max cpu.max pids.max' systemctl show -p DefaultTasksMax
memory.max:max cpu.max:max 100000 pids.max:6326 DefaultTasksMax=6326

max means no limit and cpu.max reads max 100000 (no quota, 100 ms period). pids.max is 6326 because the systemd driver gives every scope systemd's DefaultTasksMax, 15 percent of the kernel's task limit on this VM. A runaway process tree in an "unlimited" container therefore stops at a few thousand tasks instead of exhausting the host, but a few thousand per container is still enough to hurt a host running many of them. Set --pids-limit explicitly. Choosing values for real services is part of "Production best practices" in Docker in depth.

Out of memory

Give a container 64 MiB and make it hold 200 MiB. head writes 200 MiB of zeros into a pipe, and tail must keep all of it in memory because the input has no line break:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --name lab-oom --memory 64m alpine:3.22 sh -c 'head -c 200m /dev/zero | tail'
Killed
$ docker inspect -f 'OOMKilled={{.State.OOMKilled}} ExitCode={{.State.ExitCode}}' lab-oom
OOMKilled=true ExitCode=137
$ sleep 1 docker events --since 5m --until "$(date +%s)" --filter container=$(docker inspect -f '{{.Id}}' lab-oom) --filter event=oom --filter event=die --format '{{.Action}} exitCode={{.Actor.Attributes.exitCode}}' docker rm lab-oom
oom exitCode=<no value> die exitCode=137 lab-oom

Once the kernel cannot reclaim enough memory to keep the group under memory.max, the OOM killer picks a process inside that group and sends it SIGKILL, which cannot be caught. It chose tail, the process holding the memory. The shell, which is PID 1, printed Killed and exited with its pipeline's status, 137 (128 plus signal 9). The daemon saw the kernel's OOM notification and set State.OOMKilled. docker events replays the oom event followed by die with exit code 137; in production you watch the same stream live (docker events --filter event=oom) or alert on the counter below.

The cgroup itself keeps the count in memory.events, which only exists while the container does. This time stress-ng asks for 200 MiB inside a 64 MiB limit and is left running:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-hog --memory 64m --memory-swap 64m lab-stress:1 --vm 1 --vm-bytes 200M --timeout 15s
9df24b5cc312ac60b065de120fc96e9d8beabd8b56bd8f4706c79e9db278c902
$ sleep 5 cat /sys/fs/cgroup/system.slice/docker-$(docker inspect -f '{{.Id}}' lab-hog).scope/memory.events
low 0 high 0 max 51220 oom 948 oom_kill 948 oom_group_kill 0 sock_throttled 0
$ docker wait lab-hog docker inspect -f 'OOMKilled={{.State.OOMKilled}} ExitCode={{.State.ExitCode}}' lab-hog docker logs lab-hog 2>&1 | tail -3
0 OOMKilled=true ExitCode=0 stress-ng: info: [1] failed: 0 stress-ng: info: [1] metrics untrustworthy: 0 stress-ng: info: [1] successful run completed in 15.01 secs

In memory.events, max counts how often usage hit the limit, oom how often the OOM killer was invoked, and oom_kill how many processes it killed; your counts will differ. The kills landed on stress-ng's worker processes, and stress-ng started a new worker after each one and finished its 15-second run normally. The container exited 0, and OOMKilled is still true, because it means "an OOM kill happened in this container's cgroup", not "the main process died of it". That is the worker from the opening: read OOMKilled together with the exit code, and for a long-running container alert on oom_kill increasing, since a web server that loses a worker process to the OOM killer can keep running and never set a non-zero exit code.

A container reaches memory.max
Usage reaches memory.max
the kernel first tries to reclaim memory in the group
enough reclaimed
Allocation succeeds
memory.events max increases; the process sees nothing
victim is PID 1
Container exits
exit code 137, OOMKilled=true, oom and die events
victim is another process
Container keeps running
OOMKilled=true, oom_kill increases, exit code decided later by PID 1
The OOM killer only picks inside the group that hit its limit. A host-wide out-of-memory condition is a different event and can kill processes in any group.

CPU: throttled, not killed

A CPU limit never kills anything. When the group has used its quota for the current period, its threads wait until the next one. Two busy stress-ng workers under --cpus 0.5:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-cpu --cpus 0.5 lab-stress:1 --cpu 2 --timeout 60s
4737dd07178c2dab9bd3e4dc4547eb175a121b1914612033e867b9c6ffbf1ffd
$ sleep 5 CG=/sys/fs/cgroup/system.slice/docker-$(docker inspect -f '{{.Id}}' lab-cpu).scope cat "$CG/cpu.max" grep -E '^(usage_usec|nr_periods|nr_throttled|throttled_usec) ' "$CG/cpu.stat"
50000 100000 usage_usec 2309271 nr_periods 53 nr_throttled 50 throttled_usec 7951983

cpu.max is 50000 100000: 50 ms per 100 ms period. In 50 of the first 53 periods the group used its quota and was throttled, and throttled_usec adds up the waiting, about 8 seconds across both workers in 5 seconds of wall time. A latency problem with clean logs and a climbing nr_throttled is a CPU limit doing its job. systemd-cgtop gives the per-cgroup view the way top gives the per-process one; in batch mode, the columns are path, tasks, CPU percent, memory, and input and output per second:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ systemd-cgtop -b -n 2 -d 2 system.slice/docker-$(docker inspect -f '{{.Id}}' lab-cpu).scope
system.slice/docker-4737dd07178c2dab9bd3e4dc4547eb175a121b1914612033e867b9c6ffbf1ffd.scope 3 - 9.2M - - system.slice/docker-4737dd07178c2dab9bd3e4dc4547eb175a121b1914612033e867b9c6ffbf1ffd.scope 3 39.2 9.2M - -
$ docker stats --no-stream --format '{{.Name}} cpu={{.CPUPerc}} mem={{.MemUsage}} pids={{.PIDs}}' lab-cpu lab-web
lab-cpu cpu=30.88% mem=9.293MiB / 5.76GiB pids=3 lab-web cpu=0.00% mem=4.52MiB / 256MiB pids=5

The first sample has no CPU figure because cgtop needs two readings to compute a rate. Both tools take a short sample, and here they read 39 and 31 percent for lab-cpu, below its 50 percent cap. A sample that lines up badly with the 100 ms periods reads high or low, so for a precise figure use cpu.stat: usage_usec above is 2.3 seconds of CPU in about 5 seconds, close to the cap. docker stats also prints each container's memory limit: 256 MiB for lab-web, and the VM's whole memory for lab-cpu, which has none.

A fork bomb against pids.max

b(){ b|b& };b is a fork bomb written for busybox sh: each call starts two more in the background, forever. With --pids-limit 100 it fills its own group and nothing else:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-bomb --pids-limit 100 --cpus 0.5 alpine:3.22 sh -c 'b(){ b|b& };b; sleep 600'
e9f190b13150448f386057518d87879541bd0c8e52e95febb1f05c5244b6ef41
$ sleep 5 CG=/sys/fs/cgroup/system.slice/docker-$(docker inspect -f '{{.Id}}' lab-bomb).scope cat "$CG/pids.max" "$CG/pids.current" "$CG/pids.events"
100 100 max 59

pids.current sits at the limit and the max line in pids.events counts the forks the kernel refused (59 in five seconds here). Other containers and the host keep their process slots. The --cpus 0.5 on the bomb matters too: the failing fork loop still burns CPU, and without a CPU limit it would take every core it could get. docker rm -f stops it in one step: it kills PID 1, and when the first process of a PID namespace dies the kernel kills every other process in it:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker rm -f lab-web lab-hog lab-cpu lab-bomb docker rmi lab-stress:1
lab-web lab-hog lab-cpu lab-bomb Untagged: lab-stress:1 Deleted: sha256:41c216405d17e510a4151f06a5bf2a7f28909c2a8dcd1e0d974ef40c21cf4dc9

Limits like these contain a noisy or broken workload. They are not a security boundary: a container that stays under every limit can still attack the kernel through system calls and capabilities, which "Capabilities, cap-drop and no-new-privileges" takes up next.

Quick check
01A service started with --cpus 0.5 has slow responses, clean logs and no restarts. Which reading on the host confirms that the CPU limit is the cause?
Incorrect — CPU limits never kill. oom_kill counts memory OOM kills only.
Correct — Each throttled period means the group used its quota and had to wait.
Incorrect — OOMKilled is set only by memory OOM events.
Incorrect — nproc is not cgroup-aware; the lab container under a CPU limit still reports all 4 CPUs.
02A worker container with --memory 64m has stopped. docker inspect shows OOMKilled=true and ExitCode=0. Which explanation fits?
Incorrect — Reaching the limit increments the max counter in memory.events. OOMKilled needs an actual OOM event.
Incorrect — A killed main process would not exit 0, and the limit here is the container's own.
Incorrect — The flag records a real OOM kill in the container's cgroup, just not of PID 1.
Correct — As with stress-ng in the lab, the kill hit a worker and PID 1 finished with status 0.
03You start a container with --memory 256m and no other memory options. What does its memory.swap.max file contain?
Correct — MemorySwap defaults to twice --memory, and the runtime writes the difference, 256 MiB, as the swap cap.
Incorrect — Docker sets a finite default; max would need --memory-swap -1.
Incorrect — 0 is what you get with --memory-swap equal to --memory.
Incorrect — Combined memory plus swap accounting was cgroup v1. In v2 the files are separate; 536870912 is the API's MemorySwap value, not the file.

Try this

Work through “A fork bomb against pids.max” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from cgroups v2 and resource containment, keep “A fork bomb against pids.max”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related