cgroups v2 and resource containment
Memory, CPU and PID limits, OOM kills and where to read them.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-int/cgroups.tar.gz && tar -xzf cgroups.tar.gz, which creates ~/lab/cgroups/. SHA-256: 289e8645436ee1b50f59d8352704e08222a9b9e76d465f449e0aea547d458f75A batch worker has --memory 64m. After a run, docker inspect reports OOMKilled=true and ExitCode=0, and the engineer on call files it as a Docker bug, because a container that was killed for memory cannot have exited cleanly. Both values are correct. Reading them correctly needs the mechanism underneath: control groups, or cgroups, the kernel feature that counts and limits what a group of processes uses. Namespaces, from the previous lesson, decide what a container can see. cgroups decide how much memory, CPU time and how many processes it can take.
Everything runs on the main lab VM and stays inside lab-* containers. One small image, lab-stress:1, provides a load generator: its Dockerfile is in the lesson files (copy it to ~/lab/cgroups).
# lab-stress:1, a small load generator for the cgroups lesson (stress-ng from Alpine's community repo)FROM alpine:3.22RUN apk add --no-cache stress-ngENTRYPOINT ["stress-ng"]
stress-ng is maintained and packaged in Alpine. The progrium/stress image that older tutorials use is stored in the Docker schema 1 manifest format, which containerd 2.1 and later refuse to pull, so it no longer works on Docker 29.
One tree, one scope per container
Build the image in ~/lab/cgroups:
Then ask Docker and the kernel which cgroup setup this host has:
This track uses cgroup v2 only: one unified hierarchy mounted at /sys/fs/cgroup (filesystem type cgroup2fs), which current Ubuntu, Debian, Fedora and RHEL releases use by default. driver systemd means Docker asks systemd to create each container's cgroup as a transient unit called docker-<container id>.scope under system.slice, so systemd's tools see containers like any other unit. The VM has no swap (Swap: 0B), which matters for the swap section below.
Limits are files the kernel reads
Start a container with a memory limit, a CPU limit and a process limit, then read its cgroup directory on the host:
memory.current is what the group uses now (about 5 MiB for an idle nginx; yours will differ) and pids.current counts its tasks, the master and four workers. The swap line needs care. cgroup v2 has separate files for memory and swap. Docker's API still takes one combined number, MemorySwap, which defaults to twice --memory when you do not set it (536870912 above), and the runtime writes the difference into memory.swap.max. So a plain --memory 256m allows up to another 256 MiB of swap on a host that has swap. The same values also appear on the systemd unit, because systemd created the scope:
/proc/<pid>/cgroup holds a single 0:: line on cgroup v2: the process's path in the unified tree. systemd-cgls -u prints the processes in one unit. The full tree (systemd-cgls without arguments) also shows containerd and the shims inside containerd.service and dockerd inside docker.service, separate from the container scopes.
What the container sees, and what it does not
The container reads its own limit from /sys/fs/cgroup/memory.max, because its cgroup namespace shows its own cgroup as the root. /proc/meminfo and nproc are not cgroup-aware: they report the VM's 6 GB and 4 CPUs. A program that sizes a heap, a cache or a thread pool from those numbers plans for resources it will never get and is killed or throttled when it grows into them. Current Java (container support is on by default since JDK 10) and Go (GOMAXPROCS follows the CPU limit since Go 1.25) read the cgroup files; check what your runtime and your own scripts read.
Setting --memory-swap equal to --memory gives the container no swap at all:
On a host with swap, a container that reaches memory.max without that setting starts swapping instead of being killed. memory.current never goes above memory.max; the swapped pages appear in memory.swap.current, so docker stats keeps showing usage at or under the limit while the container slows down. Because the lab VM has no swap, the swap allowance here is a number in a file with no effect, and the lesson does not demonstrate swapping.
A container with no limits has no memory or CPU cap, but it does have a process cap:
max means no limit and cpu.max reads max 100000 (no quota, 100 ms period). pids.max is 6326 because the systemd driver gives every scope systemd's DefaultTasksMax, 15 percent of the kernel's task limit on this VM. A runaway process tree in an "unlimited" container therefore stops at a few thousand tasks instead of exhausting the host, but a few thousand per container is still enough to hurt a host running many of them. Set --pids-limit explicitly. Choosing values for real services is part of "Production best practices" in Docker in depth.
Out of memory
Give a container 64 MiB and make it hold 200 MiB. head writes 200 MiB of zeros into a pipe, and tail must keep all of it in memory because the input has no line break:
Once the kernel cannot reclaim enough memory to keep the group under memory.max, the OOM killer picks a process inside that group and sends it SIGKILL, which cannot be caught. It chose tail, the process holding the memory. The shell, which is PID 1, printed Killed and exited with its pipeline's status, 137 (128 plus signal 9). The daemon saw the kernel's OOM notification and set State.OOMKilled. docker events replays the oom event followed by die with exit code 137; in production you watch the same stream live (docker events --filter event=oom) or alert on the counter below.
The cgroup itself keeps the count in memory.events, which only exists while the container does. This time stress-ng asks for 200 MiB inside a 64 MiB limit and is left running:
In memory.events, max counts how often usage hit the limit, oom how often the OOM killer was invoked, and oom_kill how many processes it killed; your counts will differ. The kills landed on stress-ng's worker processes, and stress-ng started a new worker after each one and finished its 15-second run normally. The container exited 0, and OOMKilled is still true, because it means "an OOM kill happened in this container's cgroup", not "the main process died of it". That is the worker from the opening: read OOMKilled together with the exit code, and for a long-running container alert on oom_kill increasing, since a web server that loses a worker process to the OOM killer can keep running and never set a non-zero exit code.
CPU: throttled, not killed
A CPU limit never kills anything. When the group has used its quota for the current period, its threads wait until the next one. Two busy stress-ng workers under --cpus 0.5:
cpu.max is 50000 100000: 50 ms per 100 ms period. In 50 of the first 53 periods the group used its quota and was throttled, and throttled_usec adds up the waiting, about 8 seconds across both workers in 5 seconds of wall time. A latency problem with clean logs and a climbing nr_throttled is a CPU limit doing its job. systemd-cgtop gives the per-cgroup view the way top gives the per-process one; in batch mode, the columns are path, tasks, CPU percent, memory, and input and output per second:
The first sample has no CPU figure because cgtop needs two readings to compute a rate. Both tools take a short sample, and here they read 39 and 31 percent for lab-cpu, below its 50 percent cap. A sample that lines up badly with the 100 ms periods reads high or low, so for a precise figure use cpu.stat: usage_usec above is 2.3 seconds of CPU in about 5 seconds, close to the cap. docker stats also prints each container's memory limit: 256 MiB for lab-web, and the VM's whole memory for lab-cpu, which has none.
A fork bomb against pids.max
b(){ b|b& };b is a fork bomb written for busybox sh: each call starts two more in the background, forever. With --pids-limit 100 it fills its own group and nothing else:
pids.current sits at the limit and the max line in pids.events counts the forks the kernel refused (59 in five seconds here). Other containers and the host keep their process slots. The --cpus 0.5 on the bomb matters too: the failing fork loop still burns CPU, and without a CPU limit it would take every core it could get. docker rm -f stops it in one step: it kills PID 1, and when the first process of a PID namespace dies the kernel kills every other process in it:
Limits like these contain a noisy or broken workload. They are not a security boundary: a container that stays under every limit can still attack the kernel through system calls and capabilities, which "Capabilities, cap-drop and no-new-privileges" takes up next.
--cpus 0.5 has slow responses, clean logs and no restarts. Which reading on the host confirms that the CPU limit is the cause?--memory 64m has stopped. docker inspect shows OOMKilled=true and ExitCode=0. Which explanation fits?--memory 256m and no other memory options. What does its memory.swap.max file contain?Try this
Work through “A fork bomb against pids.max” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from cgroups v2 and resource containment, keep “A fork bomb against pids.max”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.