cgroup v2 resource control through systemd

CPU, memory, IO and task limits on units.

Advanced16 min · lesson 9 of 21

systemd runs every unit in a cgroup of its own, and the resource-control settings in a unit file are how you tell the kernel what that cgroup may use. This lesson applies them to running workloads and checks each effect in /sys/fs/cgroup: a slice that gives several units one shared budget, a CPU quota that throttles and a weight that does nothing without competition, a memory limit that slows a process and one that kills it, an I/O weight that no controller enforces next to a bandwidth limit that works, and limits changed on a running unit. The scheduler lesson explained cpu.weight and cpu.max from the scheduler's side and the memory lesson the OOM kill; this is the systemd interface to both.

Where the settings land

Ubuntu 26.04 and RHEL 10 mount only the unified cgroup v2 hierarchy at /sys/fs/cgroup. Each controller (cpu, memory, io, pids and others) is enabled per level of the tree: a cgroup's cgroup.subtree_control lists the controllers its children may use.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /sys/fs/cgroup/cgroup.controllers cat /sys/fs/cgroup/cgroup.subtree_control cat /sys/fs/cgroup/system.slice/cgroup.subtree_control
cpuset cpu io memory hugetlb pids rdma misc dmem cpuset cpu io memory hugetlb pids misc cpu memory pids

The first line is every controller the kernel offers. The root passes most of them down, and system.slice passes cpu, memory and pids to the services inside it, but not io. You never enable controllers by hand: systemd turns one on along the path to any unit that has a setting for it (systemd.resource-control(5)). memory and pids are on because memory and task accounting are on by default, and cpu because a packaged unit sets CPUWeight=, as the scheduler lesson found. On RHEL 10, system.slice passes only memory and pids until some unit asks for more.

systemd settings and the cgroup files they write
CPU
CPUWeight= → cpu.weight
a share, only while siblings compete
CPUQuota= → cpu.max
a ceiling per 100 ms period, even when idle
Memory
MemoryHigh= → memory.high
above it: throttled and reclaimed, never killed
MemoryMax= → memory.max
at it: reclaim, then the OOM killer in the cgroup
MemorySwapMax= → memory.swap.max
swap the cgroup may use
I/O and tasks
IOWeight= → io.weight
needs BFQ or the iocost controller
IOReadBandwidthMax= → io.max
a bandwidth ceiling on one device
TasksMax= → pids.max
processes and threads together

A slice is a node in that tree that holds other units, and settings on a slice apply to everything below it together. System services live in system.slice, logins in user.slice. This lab slice gives batch jobs one budget:

/etc/systemd/system/sd-resource.slice
[Unit]
Description=Lab batch jobs (sd-resource lesson)
[Slice]
CPUQuota=50%
MemoryHigh=96M
MemoryMax=192M
TasksMax=64

A dash in a slice name means nesting: sd-resource.slice is a child of sd.slice, which systemd creates on demand under the root slice, next to system.slice. A service joins a slice with Slice= in its unit, and systemd-run takes --slice=.

CPU: one quota for a whole slice

The workloads come from two tools that a default Ubuntu Server install does not include: stress-ng generates CPU and memory load, and fio generates disk I/O. Install them with sudo apt install stress-ng fio (on RHEL 10, sudo dnf install stress-ng fio from AppStream).

Start two CPU-bound workers in the slice, each in its own transient scope. As in the method lesson, --uid=$USER runs the workload as you, not as root, the stress-ng timeout ends it even if you forget, and the trailing & puts each scope in the background.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl daemon-reload sudo systemd-run --scope --unit=sd-resource-a --slice=sd-resource.slice --uid=$USER stress-ng --cpu 1 --timeout 5m > /dev/null 2>&1 & sudo systemd-run --scope --unit=sd-resource-b --slice=sd-resource.slice --uid=$USER stress-ng --cpu 1 --timeout 5m > /dev/null 2>&1 &
$ systemd-cgtop -b -n 2 -d 1 /sd.slice
/sd.slice 4 - 17.6M - - /sd.slice/sd-resource.slice 4 - 17.6M - - /sd.slice/sd-resource.slice/sd-resource-a.scope 2 - 10.5M - - /sd.slice/sd-resource.slice/sd-resource-b.scope 2 - 7.1M - - /sd.slice 4 50.9 17.6M - - /sd.slice/sd-resource.slice 4 51.0 17.6M - - /sd.slice/sd-resource.slice/sd-resource-b.scope 2 26.7 7.1M - - /sd.slice/sd-resource.slice/sd-resource-a.scope 2 24.3 10.5M - -

systemd-cgtop is top for cgroups; -b -n 2 -d 1 prints two samples a second apart, and /sd.slice limits it to that subtree. In batch mode it prints no header: the columns are the cgroup, its number of tasks, %CPU (100 is one full CPU), memory, and input and output per second. The first sample shows - for CPU because a rate needs two readings. In the second, the slice uses 51% of one CPU, and each worker gets about half of that, although this machine has two idle CPUs to offer: the slice's quota is shared by everything in it.

deploy@web01 · Ubuntu 26.04 LTS
$ cd /sys/fs/cgroup/sd.slice/sd-resource.slice cat cpu.max grep -E '^(nr_periods|nr_throttled|throttled_usec)' cpu.stat
50000 100000 nr_periods 63 nr_throttled 63 throttled_usec 9225117

CPUQuota=50% became cpu.max 50000 100000: 50 ms of CPU time per 100 ms period. cpu.stat shows that the slice ran out of quota in every period so far (nr_throttled equals nr_periods) and has been throttled for about 9 seconds in total, added up over both CPUs, in the 6.3 seconds of those 63 periods. That is how a quota looks when it is what limits a workload. Lift it on the running slice:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl set-property --runtime sd-resource.slice CPUQuota= systemctl cat sd-resource.slice
# /etc/systemd/system/sd-resource.slice [Unit] Description=Lab batch jobs (sd-resource lesson) [Slice] CPUQuota=50% MemoryHigh=96M MemoryMax=192M TasksMax=64 # /run/systemd/system.control/sd-resource.slice.d/50-CPUQuota.conf # This is a drop-in unit file extension, created via "systemctl set-property" # or an equivalent operation. Do not edit. [Slice] CPUQuota=

systemctl set-property changes a running unit at once. An empty value resets a setting to its default, here no quota. With --runtime the change is written to /run/systemd/system.control and lasts until the next reboot; without it, systemd writes the drop-in to /etc/systemd/system.control and it persists. systemctl cat shows the drop-in after the unit file, which is how you find a forgotten emergency change later; systemctl revert removes such drop-ins.

deploy@web01 · Ubuntu 26.04 LTS
$ systemd-cgtop -b -n 2 -d 1 /sd.slice
… /sd.slice 4 197.8 17.2M - - /sd.slice/sd-resource.slice 4 197.8 17.1M - - /sd.slice/sd-resource.slice/sd-resource-a.scope 2 99.8 10.2M - - /sd.slice/sd-resource.slice/sd-resource-b.scope 2 98.0 6.8M - -
$ sudo systemctl set-property --runtime sd-resource-a.scope CPUWeight=1000 sleep 3 systemd-cgtop -b -n 2 -d 1 /sd.slice
… /sd.slice 4 199.7 17.2M - - /sd.slice/sd-resource.slice 4 199.7 17.1M - - /sd.slice/sd-resource.slice/sd-resource-a.scope 2 100.0 10.2M - - /sd.slice/sd-resource.slice/sd-resource-b.scope 2 99.7 6.8M - -

Without the quota each worker gets a full CPU. Giving one of them CPUWeight=1000, ten times the default of 100, changed nothing: a weight divides CPU time only between siblings that compete for the same CPUs, and two workers on two CPUs do not compete. The scheduler lesson pinned two workers to one CPU and got an 83:17 split from weights of 100 and 20. Use a weight to decide who wins when there is contention and a quota when a hard ceiling is the requirement, knowing that a quota also leaves CPUs idle.

A quota has a second cost that averages hide, because it is enforced per period. A service with many threads can use its whole quota in a burst at the start of a period and then stall until the next one: with CPUQuota=200% and 16 busy threads, the 200 ms of CPU time per 100 ms period is gone after 12.5 ms, and every request that arrives in the remaining 87.5 ms waits, although the service's average CPU use is well below its limit. This is the classic latency problem with container CPU limits. Watch nr_throttled and throttled_usec in cpu.stat next to request latency, not only average CPU. CPUQuotaPeriodSec= shortens the period (1 ms to 1 s, default 100 ms), which shortens each stall. The kernel's cpu.max.burst lets a cgroup save unused quota for short bursts; systemd 259 has no setting for it. Where you do not need a hard ceiling, prefer CPUWeight=.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop sd-resource-a.scope sd-resource-b.scope

Memory: MemoryHigh is a brake, MemoryMax is a wall

Now a memory worker that wants 160 MiB, in a slice with MemoryHigh=96M and MemoryMax=192M. --vm-keep makes stress-ng keep its memory mapped and rewrite it, instead of freeing and allocating it again. The lab machine has no swap (the memory lesson removed its swap file again), so the worker's anonymous memory cannot be moved out of RAM.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --scope --unit=sd-resource-mem --slice=sd-resource.slice --uid=$USER \ stress-ng --vm 1 --vm-bytes 160M --vm-keep --timeout 5m > /var/tmp/sd-resource-mem.log 2>&1 &
$ systemctl status sd-resource.slice --lines=0
… Tasks: 3 (limit: 64) Memory: 99.7M (high: 96M, max: 192M, available: 0B, peak: 100.3M) … CGroup: /sd.slice/sd-resource.slice └─sd-resource-mem.scope ├─217003 /usr/bin/stress-ng --vm 1 --vm-bytes 160M --vm-keep --timeout 5m ├─217007 stress-ng-vm "" "" "" "" "" . └─217008 stress-ng-vm "" "" "" "" "" .

Ten seconds later the slice holds 99.7 MiB: a little above high, far below the 160 MiB the worker wants, and available: 0B says systemd sees no room left under the high limit. Tasks: 3 (limit: 64) is the slice's TasksMax=, counted across all its units. The cgroup files tell the rest:

deploy@web01 · Ubuntu 26.04 LTS
$ cd /sys/fs/cgroup/sd.slice/sd-resource.slice grep . memory.high memory.max memory.current memory.events memory.pressure
memory.high:100663296 memory.max:201326592 memory.current:104628224 memory.events:low 0 memory.events:high 3212 memory.events:max 0 memory.events:oom 0 memory.events:oom_kill 0 memory.events:oom_group_kill 0 memory.events:sock_throttled 0 memory.pressure:some avg10=42.12 avg60=9.29 avg300=2.01 total=7236601 memory.pressure:full avg10=42.12 avg60=9.29 avg300=2.01 total=7236601

memory.events high counts the times the worker was throttled and sent into direct reclaim for exceeding memory.high: 3,212 in ten seconds. memory.pressure shows the cost: in the last ten seconds, every task in the slice was stalled on memory (full) for 42% of the time. There is no OOM kill and there never will be at this limit, because the kernel documentation is explicit that going over memory.high never invokes the OOM killer. A service that lives above its MemoryHigh= is not safe but slow; raise the limit or reduce its memory use.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl set-property --runtime sd-resource.slice MemoryHigh=infinity sleep 5 cd /sys/fs/cgroup/sd.slice/sd-resource.slice grep . memory.current memory.pressure
memory.current:175435776 memory.pressure:some avg10=31.63 avg60=10.55 avg300=2.44 total=7377764 memory.pressure:full avg10=31.63 avg60=10.55 avg300=2.44 total=7377764

With the brake released, the worker reaches its full size within seconds (167 MiB for the slice). The pressure averages decay slowly over their windows, so look at total, the microseconds stalled since the cgroup was created: it grew by about 0.14 s in these five seconds, against 7.2 s in the ten seconds before. systemd.resource-control(5) recommends MemoryHigh= as the main control and MemoryMax= as the last line of defence. Set the high limit below the maximum so that pressure shows up in memory.events and memory.pressure before anything is killed. Reaching MemoryMax= means reclaim and then the OOM killer inside the cgroup, followed by the unit's OOMPolicy=, as the memory lesson showed. Stopping a slice stops every unit in it:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop sd-resource.slice pgrep -c stress-ng
0

I/O: weights need a controller that honours them

IOWeight= writes io.weight (1 to 10000, default 100), which asks for a proportional share of disk time. A weight is only a number until a controller acts on it, and the kernel has two that do: the BFQ I/O scheduler, which has its own io.bfq.weight file that systemd sets from the same setting, and the iocost controller, configured in io.cost.qos on the root cgroup and disabled by default.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /sys/block/vda/queue/scheduler wc -l < /sys/fs/cgroup/io.cost.qos
none [mq-deadline] 0

This disk uses the mq-deadline scheduler, with none the only alternative offered (Ubuntu builds BFQ as a module, which is not loaded); io.cost.qos is empty, so iocost is off. IOWeight= on any unit here has no effect at all, even when two units fight over the disk. RHEL 10 lists kyber and bfq as well, but also defaults to mq-deadline. Making weights work means switching the device to BFQ (a udev rule makes it permanent) or configuring iocost, and both change I/O behaviour for everything on the device, so measure before and after. A bandwidth limit, io.max, works with any scheduler. First the disk without a limit, reading a 64 MiB file with fio and direct I/O so the page cache does not answer instead of the disk:

deploy@web01 · Ubuntu 26.04 LTS
$ fio --name=read --filename=/var/tmp/sd-resource.dat --rw=read --bs=1M --size=64M --direct=1
… read: IOPS=1882, BW=1882MiB/s (1974MB/s)(64.0MiB/34msec) …

Then the same read in a scope with IOReadBandwidthMax=, which takes a device, or a path whose device systemd looks up, and a rate in bytes per second, where 8M means 8,000,000:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --scope --unit=sd-resource-io --uid=$USER -p IOReadBandwidthMax="/var/tmp 8M" \ fio --name=read --filename=/var/tmp/sd-resource.dat --rw=read --bs=1M --size=64M --direct=1 > /var/tmp/sd-resource-io.log 2>&1 &
$ cat /sys/fs/cgroup/system.slice/sd-resource-io.scope/io.max lsblk -dno MAJ:MIN,NAME /dev/vda
253:0 rbps=8000000 wbps=max riops=max wiops=max 253:0 vda
$ cat /var/tmp/sd-resource-io.log
Running as unit: sd-resource-io.scope; invocation ID: d343c27630dc4574af434cdea9649397 … read: IOPS=7, BW=7807KiB/s (7995kB/s)(64.0MiB/8394msec) …

Unlimited, the 64 MiB read took 34 ms (1,882 MiB/s). That number says nothing about the disk: direct I/O bypasses only the guest's page cache, and on this VM the host's own file cache holding the disk image most likely answered. A benchmark inside a VM measures the whole virtual stack, not the SSD underneath; all that matters here is that it is far faster than the limit. systemd resolved /var/tmp to the whole disk vda, device 253:0, and wrote rbps=8000000 to the scope's io.max; the same read then ran at 7,995 kB/s and took 8.4 seconds. Like a CPU quota, io.max is not work-conserving: the limit applies even when the device is idle.

Try this

Put a wall under a running workload. Start the memory worker again with the same systemd-run command (the runtime MemoryHigh=infinity from earlier still applies to the slice until the next reboot), wait five seconds, and lower the slice's hard limit below what it holds: sudo systemctl set-property --runtime sd-resource.slice MemoryMax=128M. Predict what systemctl status sd-resource-mem.scope shows, and which cgroup the kernel names. Expect the scope to be failed (Result: oom-kill), because the default OOMPolicy=stop applies to scopes as well, and the oom-kill: line in journalctl -k to name the slice as oom_memcg (the limit that was hit) and the scope as task_memcg (where the victim lived). Finish with sudo systemctl stop sd-resource.slice, sudo systemctl revert sd-resource.slice, then delete the slice file and /var/tmp/sd-resource.dat and run sudo systemctl daemon-reload.

Takeaway

Choose the setting by the behaviour you want under pressure: CPUWeight= and a MemoryHigh= below MemoryMax= for workloads that should yield and slow down, quotas and MemoryMax= for hard ceilings, and io.max limits instead of IOWeight= unless the disk runs BFQ or iocost. Then read the cgroup's files to confirm the kernel is doing what the unit file says.

Quick check
01A service's memory.current sits just above its MemoryHigh all day, memory.events shows high climbing by thousands an hour, and oom_kill is 0. Users report that it is slow. What is going on?
Incorrect — Above memory.high the kernel throttles the cgroup and forces reclaim on every allocation, which is a performance problem, not a target.
Correct — The high counter is the number of throttling events, and memory.pressure would show the time the service loses to them.
Incorrect — memory.high never invokes the OOM killer, and oom_kill 0 confirms that nothing was killed.
Incorrect — The high counter records throttling. Allocations above memory.high succeed, only slowly.
02You set IOWeight=10 on a backup service and IOWeight=1000 on a database. During backups the database's I/O latency does not change. The disk's scheduler file shows [none] and io.cost.qos is empty. Why?
Incorrect — systemd writes io.weight immediately. The value is present in the cgroup; nothing acts on it.
Incorrect — A ratio of 1:100 is a large difference in any proportional controller. The problem is that there is none.
Correct — Proportional I/O needs BFQ or iocost. With none and iocost disabled, only io.max limits apply.
Incorrect — cgroup limits apply to every process in the cgroup, whoever owns it.
03During an incident you ran sudo systemctl set-property api.service MemoryMax=1G. Weeks and several reboots later you edit api.service to MemoryMax=4G and run daemon-reload, but the service still gets 1 GiB. Why?
Correct — Without --runtime the change persists as a drop-in, and drop-ins are applied after the main file. systemctl cat shows it and systemctl revert removes it.
Incorrect — cgroups live in memory and are recreated at every boot from the unit configuration.
Incorrect — daemon-reload re-reads unit files and drop-ins; the old value comes from a file, not a cache.
Incorrect — Limits can be raised at any time, with set-property or an edited unit and a reload.

Related