cgroup v2 resource control through systemd
CPU, memory, IO and task limits on units.
systemd runs every unit in a cgroup of its own, and the resource-control settings in a unit file are how you tell the kernel what that cgroup may use. This lesson applies them to running workloads and checks each effect in /sys/fs/cgroup: a slice that gives several units one shared budget, a CPU quota that throttles and a weight that does nothing without competition, a memory limit that slows a process and one that kills it, an I/O weight that no controller enforces next to a bandwidth limit that works, and limits changed on a running unit. The scheduler lesson explained cpu.weight and cpu.max from the scheduler's side and the memory lesson the OOM kill; this is the systemd interface to both.
Where the settings land
Ubuntu 26.04 and RHEL 10 mount only the unified cgroup v2 hierarchy at /sys/fs/cgroup. Each controller (cpu, memory, io, pids and others) is enabled per level of the tree: a cgroup's cgroup.subtree_control lists the controllers its children may use.
The first line is every controller the kernel offers. The root passes most of them down, and system.slice passes cpu, memory and pids to the services inside it, but not io. You never enable controllers by hand: systemd turns one on along the path to any unit that has a setting for it (systemd.resource-control(5)). memory and pids are on because memory and task accounting are on by default, and cpu because a packaged unit sets CPUWeight=, as the scheduler lesson found. On RHEL 10, system.slice passes only memory and pids until some unit asks for more.
A slice is a node in that tree that holds other units, and settings on a slice apply to everything below it together. System services live in system.slice, logins in user.slice. This lab slice gives batch jobs one budget:
[Unit]Description=Lab batch jobs (sd-resource lesson)[Slice]CPUQuota=50%MemoryHigh=96MMemoryMax=192MTasksMax=64
A dash in a slice name means nesting: sd-resource.slice is a child of sd.slice, which systemd creates on demand under the root slice, next to system.slice. A service joins a slice with Slice= in its unit, and systemd-run takes --slice=.
CPU: one quota for a whole slice
The workloads come from two tools that a default Ubuntu Server install does not include: stress-ng generates CPU and memory load, and fio generates disk I/O. Install them with sudo apt install stress-ng fio (on RHEL 10, sudo dnf install stress-ng fio from AppStream).
Start two CPU-bound workers in the slice, each in its own transient scope. As in the method lesson, --uid=$USER runs the workload as you, not as root, the stress-ng timeout ends it even if you forget, and the trailing & puts each scope in the background.
systemd-cgtop is top for cgroups; -b -n 2 -d 1 prints two samples a second apart, and /sd.slice limits it to that subtree. In batch mode it prints no header: the columns are the cgroup, its number of tasks, %CPU (100 is one full CPU), memory, and input and output per second. The first sample shows - for CPU because a rate needs two readings. In the second, the slice uses 51% of one CPU, and each worker gets about half of that, although this machine has two idle CPUs to offer: the slice's quota is shared by everything in it.
CPUQuota=50% became cpu.max 50000 100000: 50 ms of CPU time per 100 ms period. cpu.stat shows that the slice ran out of quota in every period so far (nr_throttled equals nr_periods) and has been throttled for about 9 seconds in total, added up over both CPUs, in the 6.3 seconds of those 63 periods. That is how a quota looks when it is what limits a workload. Lift it on the running slice:
systemctl set-property changes a running unit at once. An empty value resets a setting to its default, here no quota. With --runtime the change is written to /run/systemd/system.control and lasts until the next reboot; without it, systemd writes the drop-in to /etc/systemd/system.control and it persists. systemctl cat shows the drop-in after the unit file, which is how you find a forgotten emergency change later; systemctl revert removes such drop-ins.
Without the quota each worker gets a full CPU. Giving one of them CPUWeight=1000, ten times the default of 100, changed nothing: a weight divides CPU time only between siblings that compete for the same CPUs, and two workers on two CPUs do not compete. The scheduler lesson pinned two workers to one CPU and got an 83:17 split from weights of 100 and 20. Use a weight to decide who wins when there is contention and a quota when a hard ceiling is the requirement, knowing that a quota also leaves CPUs idle.
A quota has a second cost that averages hide, because it is enforced per period. A service with many threads can use its whole quota in a burst at the start of a period and then stall until the next one: with CPUQuota=200% and 16 busy threads, the 200 ms of CPU time per 100 ms period is gone after 12.5 ms, and every request that arrives in the remaining 87.5 ms waits, although the service's average CPU use is well below its limit. This is the classic latency problem with container CPU limits. Watch nr_throttled and throttled_usec in cpu.stat next to request latency, not only average CPU. CPUQuotaPeriodSec= shortens the period (1 ms to 1 s, default 100 ms), which shortens each stall. The kernel's cpu.max.burst lets a cgroup save unused quota for short bursts; systemd 259 has no setting for it. Where you do not need a hard ceiling, prefer CPUWeight=.
Memory: MemoryHigh is a brake, MemoryMax is a wall
Now a memory worker that wants 160 MiB, in a slice with MemoryHigh=96M and MemoryMax=192M. --vm-keep makes stress-ng keep its memory mapped and rewrite it, instead of freeing and allocating it again. The lab machine has no swap (the memory lesson removed its swap file again), so the worker's anonymous memory cannot be moved out of RAM.
Ten seconds later the slice holds 99.7 MiB: a little above high, far below the 160 MiB the worker wants, and available: 0B says systemd sees no room left under the high limit. Tasks: 3 (limit: 64) is the slice's TasksMax=, counted across all its units. The cgroup files tell the rest:
memory.events high counts the times the worker was throttled and sent into direct reclaim for exceeding memory.high: 3,212 in ten seconds. memory.pressure shows the cost: in the last ten seconds, every task in the slice was stalled on memory (full) for 42% of the time. There is no OOM kill and there never will be at this limit, because the kernel documentation is explicit that going over memory.high never invokes the OOM killer. A service that lives above its MemoryHigh= is not safe but slow; raise the limit or reduce its memory use.
With the brake released, the worker reaches its full size within seconds (167 MiB for the slice). The pressure averages decay slowly over their windows, so look at total, the microseconds stalled since the cgroup was created: it grew by about 0.14 s in these five seconds, against 7.2 s in the ten seconds before. systemd.resource-control(5) recommends MemoryHigh= as the main control and MemoryMax= as the last line of defence. Set the high limit below the maximum so that pressure shows up in memory.events and memory.pressure before anything is killed. Reaching MemoryMax= means reclaim and then the OOM killer inside the cgroup, followed by the unit's OOMPolicy=, as the memory lesson showed. Stopping a slice stops every unit in it:
I/O: weights need a controller that honours them
IOWeight= writes io.weight (1 to 10000, default 100), which asks for a proportional share of disk time. A weight is only a number until a controller acts on it, and the kernel has two that do: the BFQ I/O scheduler, which has its own io.bfq.weight file that systemd sets from the same setting, and the iocost controller, configured in io.cost.qos on the root cgroup and disabled by default.
This disk uses the mq-deadline scheduler, with none the only alternative offered (Ubuntu builds BFQ as a module, which is not loaded); io.cost.qos is empty, so iocost is off. IOWeight= on any unit here has no effect at all, even when two units fight over the disk. RHEL 10 lists kyber and bfq as well, but also defaults to mq-deadline. Making weights work means switching the device to BFQ (a udev rule makes it permanent) or configuring iocost, and both change I/O behaviour for everything on the device, so measure before and after. A bandwidth limit, io.max, works with any scheduler. First the disk without a limit, reading a 64 MiB file with fio and direct I/O so the page cache does not answer instead of the disk:
Then the same read in a scope with IOReadBandwidthMax=, which takes a device, or a path whose device systemd looks up, and a rate in bytes per second, where 8M means 8,000,000:
Unlimited, the 64 MiB read took 34 ms (1,882 MiB/s). That number says nothing about the disk: direct I/O bypasses only the guest's page cache, and on this VM the host's own file cache holding the disk image most likely answered. A benchmark inside a VM measures the whole virtual stack, not the SSD underneath; all that matters here is that it is far faster than the limit. systemd resolved /var/tmp to the whole disk vda, device 253:0, and wrote rbps=8000000 to the scope's io.max; the same read then ran at 7,995 kB/s and took 8.4 seconds. Like a CPU quota, io.max is not work-conserving: the limit applies even when the device is idle.
Try this
Put a wall under a running workload. Start the memory worker again with the same systemd-run command (the runtime MemoryHigh=infinity from earlier still applies to the slice until the next reboot), wait five seconds, and lower the slice's hard limit below what it holds: sudo systemctl set-property --runtime sd-resource.slice MemoryMax=128M. Predict what systemctl status sd-resource-mem.scope shows, and which cgroup the kernel names. Expect the scope to be failed (Result: oom-kill), because the default OOMPolicy=stop applies to scopes as well, and the oom-kill: line in journalctl -k to name the slice as oom_memcg (the limit that was hit) and the scope as task_memcg (where the victim lived). Finish with sudo systemctl stop sd-resource.slice, sudo systemctl revert sd-resource.slice, then delete the slice file and /var/tmp/sd-resource.dat and run sudo systemctl daemon-reload.
Takeaway
Choose the setting by the behaviour you want under pressure: CPUWeight= and a MemoryHigh= below MemoryMax= for workloads that should yield and slow down, quotas and MemoryMax= for hard ceilings, and io.max limits instead of IOWeight= unless the disk runs BFQ or iocost. Then read the cgroup's files to confirm the kernel is doing what the unit file says.