Virtual memory, page cache and the OOM killer

RSS vs PSS, reclaim, swap and OOM.

Advanced18 min · lesson 8 of 21

Memory questions on a server rarely have the answer the first number suggests: RSS that adds up to more than the machine has, an allocation that succeeds on a full machine, a service killed while free still shows gigabytes available. This lesson follows memory from a process's address space to RAM, to disk and back: what a process really uses, what the kernel promises, how the page cache fills and is written back, minor and major page faults, a cgroup that swaps, and out-of-memory kills inside cgroup limits with the kernel's report and systemd's response. Linux essentials covered free and "available"; the /proc lesson introduced smaps_rollup.

Address space and resident memory

A process sees virtual memory: a private range of addresses mapped to files, shared libraries and anonymous memory (heap and stack, backed by no file). Mapping a range costs almost nothing. The kernel assigns a physical page only when the process first touches an address in it, a scheme called demand paging. This program reserves 1 GiB, writes to 200 MiB of it and then forks, so parent and child share those pages. Save it as /var/tmp/k-mem-share and make it executable (chmod +x).

/var/tmp/k-mem-share
#!/usr/bin/python3
# k-mem-share: reserves 1 GiB of address space, writes to 200 MiB of it, then forks.
import mmap, os, time
region = mmap.mmap(-1, 1 << 30, flags=mmap.MAP_PRIVATE) # 1 GiB, nothing touched yet
for offset in range(0, 200 << 20, mmap.PAGESIZE): # one write per 4 KiB page
region[offset] = 1
os.fork() # parent and child share the pages
time.sleep(3600)
deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --unit=k-mem-share --uid=$USER -p MemoryMax=512M /var/tmp/k-mem-share
Running as unit: k-mem-share.service; invocation ID: b906d5e13c5f49459b83304728f06b8f
$ ps -o pid,ppid,vsz,rss,comm -C k-mem-share
PID PPID VSZ RSS COMMAND 260117 1 1068844 216048 k-mem-share 260135 260117 1068844 212844 k-mem-share

ps reports both columns in KiB. VSZ (virtual size) is the address space, about 1 GiB plus Python's own mappings. RSS (resident set size) counts pages actually in RAM, about 210 MiB each, so the two seem to use over 420 MiB.

deploy@web01 · Ubuntu 26.04 LTS
$ grep -E '^(Rss|Pss|Pss_Anon|Shared_Dirty|Private_Dirty):' /proc/$(pgrep -ox k-mem-share)/smaps_rollup /proc/$(pgrep -nx k-mem-share)/smaps_rollup
/proc/260117/smaps_rollup:Rss: 216048 kB /proc/260117/smaps_rollup:Pss: 106465 kB /proc/260117/smaps_rollup:Pss_Anon: 105122 kB /proc/260117/smaps_rollup:Shared_Dirty: 210004 kB /proc/260117/smaps_rollup:Private_Dirty: 120 kB /proc/260135/smaps_rollup:Rss: 212844 kB /proc/260135/smaps_rollup:Pss: 105605 kB /proc/260135/smaps_rollup:Pss_Anon: 105130 kB /proc/260135/smaps_rollup:Shared_Dirty: 210004 kB /proc/260135/smaps_rollup:Private_Dirty: 128 kB

They do not. After fork() the child gets a copy of the parent's page tables, not of its memory: both point at the same physical pages, marked copy-on-write, and a page is copied only when one side writes to it. Shared_Dirty (about 205 MiB) is that shared memory. Pss (proportional set size) charges each shared page to its sharers in equal parts, about 104 MiB each, so the PSS values of a group of processes add up to what the group really occupies.

deploy@web01 · Ubuntu 26.04 LTS
$ systemctl status k-mem-share --lines=0
… Tasks: 2 (limit: 4107) Memory: 206.6M (max: 512M, available: 305.3M, peak: 206.6M) …

The cgroup counts every page once, against the cgroup that first used it: the service uses 206.6 MiB. To size a service, read its cgroup (systemctl status, or memory.current as later in this lesson); to compare processes, use PSS; never add up RSS.

Overcommit: promises and pages

The 1 GiB reservation cost no RAM, but the kernel does account for it. Every private writable mapping is a promise that pages will be available when touched, and Committed_AS in /proc/meminfo totals those promises.

deploy@web01 · Ubuntu 26.04 LTS
$ sysctl vm.overcommit_memory grep -E '^(MemTotal|CommitLimit|Committed_AS):' /proc/meminfo
vm.overcommit_memory = 0 MemTotal: 3989352 kB CommitLimit: 1994676 kB Committed_AS: 2587248 kB

Committed_AS is 2.6 GB on a 3.8 GB machine because each of the two processes holds its own 1 GiB promise: after a fork, either could write to every page. It is above CommitLimit (swap plus 50% of RAM, vm.overcommit_ratio), and nothing complains, because that limit is enforced only in mode 2. The default on Ubuntu and RHEL, vm.overcommit_memory=0, is a heuristic.

deploy@web01 · Ubuntu 26.04 LTS
$ python3 -c 'import mmap; m = mmap.mmap(-1, 3 << 30, flags=mmap.MAP_PRIVATE); print("reserved 3 GiB")'
reserved 3 GiB
$ python3 -c 'import mmap; m = mmap.mmap(-1, 8 << 30, flags=mmap.MAP_PRIVATE); print("reserved 8 GiB")'
Traceback (most recent call last): … OSError: [Errno 12] Cannot allocate memory
$ sudo systemctl stop k-mem-share

The heuristic compares a single request with RAM plus swap, not with what is free or already promised: 3 GiB, less than the 3.8 GiB of RAM, is granted, and 8 GiB is refused with ENOMEM. So an out-of-memory failure rarely happens at malloc(); it happens later, when pages are touched and none can be found, and the answer then is the OOM killer. Mode 1 grants everything. Mode 2 fails allocations at CommitLimit instead, which programs must handle, and it is no switch to flip on a running server: here Committed_AS is already above CommitLimit, so in mode 2 the next allocation, even the fork() of an SSH login, would fail.

The page cache and dirty pages

File data passes through the page cache, the RAM the kernel keeps copies of file pages in. A read fills it; a write goes into it and marks the pages dirty, to be written to disk later. fincore (util-linux) shows how much of a file is in the cache and how much of that is dirty.

deploy@web01 · Ubuntu 26.04 LTS
$ head -c 64M /dev/urandom > /var/tmp/k-mem.dat fincore -o RES,DIRTY,FILE /var/tmp/k-mem.dat
RES DIRTY FILE 64M 64M /var/tmp/k-mem.dat
$ grep -E '^(MemTotal|MemFree|MemAvailable|Cached|Dirty|Writeback|Active\(anon\)|Inactive\(anon\)|Active\(file\)|Inactive\(file\)):' /proc/meminfo
MemTotal: 3989352 kB MemFree: 2636872 kB MemAvailable: 3451744 kB Cached: 669784 kB Active(anon): 113692 kB Inactive(anon): 30352 kB Active(file): 146012 kB Inactive(file): 554272 kB Dirty: 67000 kB Writeback: 0 kB

head finished as soon as its 64 MiB were in memory; all of it is resident and dirty, and Dirty in /proc/meminfo includes it. MemAvailable, as the essentials lesson explained, also counts cache that can be dropped, which is why a healthy server shows little free memory. When memory is needed the kernel reclaims pages, dropping or writing out their contents so they can be reused. It ages anonymous and file pages separately from active (recently used) to inactive and takes the oldest first; both lab kernels do this with the multi-gen LRU, and /proc/meminfo shows its Active and Inactive as an approximation. Most memory in use here is file cache, much of it inactive: cheap to drop.

deploy@web01 · Ubuntu 26.04 LTS
$ sysctl vm.dirty_background_ratio vm.dirty_ratio vm.dirty_expire_centisecs vm.dirty_writeback_centisecs
vm.dirty_background_ratio = 10 vm.dirty_ratio = 20 vm.dirty_expire_centisecs = 3000 vm.dirty_writeback_centisecs = 500

Flusher threads wake every 5 s (dirty_writeback_centisecs 500) and write pages dirty for more than 30 s (dirty_expire_centisecs 3000). They also start once dirty pages exceed 10% of the memory that could hold them (dirty_background_ratio; free and reclaimable pages, not total RAM), and at 20% (dirty_ratio) a writing process has to write out dirty data itself, which slows it down. The RHEL 10 kernel has the same defaults, with a caveat at the end of this lesson.

deploy@web01 · Ubuntu 26.04 LTS
$ sync fincore -o RES,DIRTY,FILE /var/tmp/k-mem.dat
RES DIRTY FILE 64M 0B /var/tmp/k-mem.dat

sync forced the writeback: the file is still cached but clean, so the kernel can drop those pages at no cost. Until writeback, the only copy of those 64 MiB was in RAM; a program that must know its data is on disk calls fsync(), as the filesystem lesson shows.

Page faults, reclaim and swap

A page fault happens when a process touches an address that has no page mapped yet. A minor fault is resolved from memory: the page is already in the page cache, or a fresh zeroed page will do. A major fault needs I/O, a read from disk or swap, and the process waits for it. To see the difference, the lab first empties the page cache.

Dropping caches is a lab tool only
Writing 1 to /proc/sys/vm/drop_caches discards every clean page-cache page on the machine. On a production server it fixes nothing (the kernel reclaims cache by itself when memory is needed) and causes a burst of disk reads as everything is read back. Use it only to make a cold-cache test repeatable on a lab machine.
/var/tmp/k-mem-faults
#!/usr/bin/python3
# k-mem-faults: maps a file, reads one byte of every page, twice, and prints how many
# page faults each pass took. With the argument random it turns readahead off.
import mmap, resource, sys
def faults():
usage = resource.getrusage(resource.RUSAGE_SELF)
return usage.ru_minflt, usage.ru_majflt
with open(sys.argv[1], 'rb') as f:
for n in (1, 2):
m = mmap.mmap(f.fileno(), 0, prot=mmap.PROT_READ)
if sys.argv[2:] == ['random']:
m.madvise(mmap.MADV_RANDOM)
before = faults()
for offset in range(0, len(m), mmap.PAGESIZE):
m[offset]
after = faults()
print(f'pass {n}: {after[0] - before[0]} minor, {after[1] - before[1]} major')
m.close()

Save it as /var/tmp/k-mem-faults and make it executable in the same way.

deploy@web01 · Ubuntu 26.04 LTS
$ echo 1 | sudo tee /proc/sys/vm/drop_caches fincore -o RES,FILE /var/tmp/k-mem.dat
1 RES FILE 0B /var/tmp/k-mem.dat
$ /var/tmp/k-mem-faults /var/tmp/k-mem.dat
pass 1: 435 minor, 1 major pass 2: 436 minor, 0 major
$ fincore -o RES,FILE /var/tmp/k-mem.dat
RES FILE 64M /var/tmp/k-mem.dat

The cache held none of the file, yet reading all 16,384 pages cost one major fault. The kernel saw sequential access and read ahead, so the later faults found their pages already in memory. The few hundred minor faults per pass are far fewer than pages, because each fault maps a batch of cached pages around the address (fault-around). Turn readahead off, as a random-access workload effectively does:

deploy@web01 · Ubuntu 26.04 LTS
$ echo 1 | sudo tee /proc/sys/vm/drop_caches > /dev/null /var/tmp/k-mem-faults /var/tmp/k-mem.dat random
pass 1: 0 minor, 16384 major pass 2: 1024 minor, 0 major

Now every page is a major fault: 16,384 separate disk reads, each one a wait. On the second pass the file is cached and every fault is minor: 1,024 of them, one per 16 pages, because fault-around maps the cached pages in the 64 KiB around each fault. This is how a database whose working set does not fit in the cache becomes disk-bound. pidstat -r shows the rate per process (majflt/s); /proc/vmstat keeps the system totals:

deploy@web01 · Ubuntu 26.04 LTS
$ grep -E '^(pgmajfault|pgscan_kswapd|pgscan_direct|allocstall_normal) ' /proc/vmstat
allocstall_normal 0 pgmajfault 51830 pgscan_kswapd 13561 pgscan_direct 0

When free memory falls below a low watermark, the kswapd kernel thread reclaims pages in the background (pgscan_kswapd). If an allocation still finds too little, the allocating task reclaims pages itself and stalls: direct reclaim (pgscan_direct, allocstall_*), part of what memory pressure (PSI) measures. Here kswapd has had some work since boot (other workloads on this lab machine), and no allocation has ever waited for direct reclaim. Clean file pages are dropped, dirty ones written first, and anonymous pages can only go to swap.

deploy@web01 · Ubuntu 26.04 LTS
$ swapon --show cat /proc/sys/vm/swappiness
60

Neither lab machine has swap. vm.swappiness (0 to 200, default 60) is the relative cost the kernel assumes for swapping anonymous pages against re-reading file pages: 100 means equal cost, lower values protect anonymous memory. Without swap, pressure can only evict file pages, including the code of running programs, which must then be read back. To see swap at work, add a swap file on the lab machine:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo fallocate -l 512M /var/tmp/k-mem.swap sudo chmod 600 /var/tmp/k-mem.swap sudo mkswap /var/tmp/k-mem.swap sudo swapon /var/tmp/k-mem.swap swapon --show
Setting up swapspace version 1, size = 512 MiB (536866816 bytes) no label, UUID=d53cfed9-0363-490f-936a-4ecdea665c6a NAME TYPE SIZE USED PRIO /var/tmp/k-mem.swap file 512M 0B -1

The next program allocates and fills a given number of MiB and waits. Save it as /var/tmp/k-mem-hog and make it executable. It runs in a scope limited to 64 MiB and asks for 128 MiB.

/var/tmp/k-mem-hog
#!/usr/bin/python3
# k-mem-hog: allocates and fills the number of MiB given as its argument, then waits.
import sys, time
size = int(sys.argv[1])
ballast = b'x' * (size << 20)
print(f'holding {size} MiB', flush=True)
time.sleep(600)
deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --scope --unit=k-mem-swap --uid=$USER -p MemoryMax=64M \ /var/tmp/k-mem-hog 128 > /var/tmp/k-mem-swap.log 2>&1 &
$ cat /var/tmp/k-mem-swap.log
Running as unit: k-mem-swap.scope; invocation ID: 224baba2f46b4aec9f579bd8a747727c holding 128 MiB
$ cd /sys/fs/cgroup/system.slice/k-mem-swap.scope grep . memory.max memory.current memory.swap.current memory.pressure
memory.max:67108864 memory.current:66916352 memory.swap.current:74788864 memory.pressure:some avg10=0.00 avg60=0.00 avg300=0.00 total=40009 memory.pressure:full avg10=0.00 avg60=0.00 avg300=0.00 total=40009
$ grep -E '^(anon|file|pgscan_direct|pgsteal_direct|pswpout|pgmajfault) ' /sys/fs/cgroup/system.slice/k-mem-swap.scope/memory.stat
anon 63369216 file 0 pswpout 18163 pgscan_direct 65435 pgsteal_direct 17496 pgmajfault 106

The program got its 128 MiB and survived. memory.current sits at the 64 MiB limit, 71 MiB are in swap and 60 MiB of anonymous memory in RAM: the limit caps RAM, and with swap available, reclaim inside the cgroup swapped pages out (pswpout, in pages) instead of killing anything. The allocating task did that reclaim itself (pgscan_direct), and memory.pressure records 40 ms of stall. On a host with swap, MemoryMax= therefore turns an out-of-memory condition into a slower program unless MemorySwapMax=0 is set too. Note the OOM score choom reports; the next section explains it.

deploy@web01 · Ubuntu 26.04 LTS
$ choom -p $(pgrep -x k-mem-hog)
pid 260638's current OOM score: 687 pid 260638's current OOM score adjust value: 0
$ sudo systemctl stop k-mem-swap.scope

Remove the swap file again, so that the rest of this lesson and the course run without swap, as both lab machines do:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo swapoff /var/tmp/k-mem.swap sudo rm /var/tmp/k-mem.swap swapon --show

When memory runs out: the OOM killer

The kernel runs out of memory in one of two places: globally, when the whole machine has nothing left to reclaim, or in a cgroup that reaches memory.max when reclaim inside it cannot bring usage down. Either way the out-of-memory (OOM) killer picks one task among the candidates (all tasks, or the cgroup's) and kills it with SIGKILL. The lab only causes cgroup OOMs, so the VM itself never runs short.

What counts toward the limit is more than the heap. memory.current, the number memory.max is compared with, includes the page cache the service's own reads and writes bring in and the kernel memory charged to it, such as dentry and inode slab objects, socket buffers and page tables; memory.stat in the cgroup breaks it down (anon, file, slab, sock, pagetables). A file-heavy service therefore meets its limit first by reclaiming its own cache, long before any OOM, and its read latency rises instead. Read anon and file in memory.stat before you size a limit from memory.peak.

What the kernel does when a task needs a page
1Free pages above the low watermark
the page is handed out at once
2kswapd reclaims in the background
woken below the low watermark
3Direct reclaim by the allocating task
the task stalls; memory PSI counts it
4Drop clean cache, write dirty, swap anon
anonymous pages move only if there is swap
5OOM killer
nothing reclaimable left: kill one task
In a cgroup, reaching memory.max skips kswapd: the allocating task reclaims within the cgroup, then the cgroup OOM killer acts.

This service runs the hog through sh, so the main process is the shell and the hog is its child. MemorySwapMax=0 keeps it out of swap on a host that has some.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --unit=k-mem-oom --uid=$USER -p MemoryMax=64M -p MemorySwapMax=0 \ sh -c '/var/tmp/k-mem-hog 128; exec sleep 600'
Running as unit: k-mem-oom.service; invocation ID: 450e6ceca6d94868acc148f86cef7475
$ systemctl status k-mem-oom
× k-mem-oom.service - [systemd-run] /usr/bin/sh -c "/var/tmp/k-mem-hog 128; exec sleep 600" … Active: failed (Result: oom-kill) since Sun 2026-09-27 09:45:37 UTC; 115ms ago … Process: 260817 ExecStart=/usr/bin/sh -c /var/tmp/k-mem-hog 128; exec sleep 600 (code=killed, signal=TERM) … Sep 27 09:45:37 web01 systemd[1]: Started k-mem-oom.service - [systemd-run] /usr/bin/sh -c "/var/tmp/k-mem-hog 128; exec sleep 600". Sep 27 09:45:37 web01 sh[260817]: Killed Sep 27 09:45:37 web01 systemd[1]: k-mem-oom.service: The kernel OOM killer killed some processes in this unit. Sep 27 09:45:37 web01 systemd[1]: k-mem-oom.service: Failed with result 'oom-kill'.

The hog was killed (sh reports its child as Killed), and systemd logged the kernel OOM kill. Then systemd stopped the whole unit, sending SIGTERM to the healthy shell, and marked it failed with the result oom-kill. That is OOMPolicy=stop, the default (DefaultOOMPolicy= in systemd-system.conf(5)) for units that do not delegate their cgroup; with Restart=on-failure the unit would come back. Units with Delegate=yes, such as container managers and user@.service, default to continue (systemd.service(5)), as the namespaces lesson shows. The kernel's own account:

deploy@web01 · Ubuntu 26.04 LTS
$ journalctl -k -b -o cat | grep -E 'invoked oom-killer|^memory: usage|^swap: usage|^oom-kill:|out of memory: Killed' | tail -5
k-mem-hog invoked oom-killer: gfp_mask=0xcc0(GFP_KERNEL), order=0, oom_score_adj=0 memory: usage 65536kB, limit 65536kB, failcnt 37 swap: usage 0kB, limit 0kB, failcnt 0 oom-kill:constraint=CONSTRAINT_MEMCG,nodemask=(null),cpuset=system.slice,mems_allowed=0,oom_memcg=/system.slice/k-mem-oom.service,task_memcg=/system.slice/k-mem-oom.service,task=k-mem-hog,pid=260829,uid=1001 Memory cgroup out of memory: Killed process 260829 (k-mem-hog) total-vm:151344kB, anon-rss:65116kB, file-rss:5696kB, shmem-rss:0kB, UID:1001 pgtables:196kB oom_score_adj:0

These are five lines of a longer report with a stack trace, memory statistics and a table of candidate tasks. The first names the task whose allocation failed, not necessarily the victim. memory: usage equals the limit and failcnt counts how often it was hit; swap: limit 0kB is MemorySwapMax=0. CONSTRAINT_MEMCG in the oom-kill: line marks a cgroup OOM, with oom_memcg the cgroup whose limit was reached and task_memcg the victim's; a global OOM says CONSTRAINT_NONE. The last line names the victim and its 64 MiB of anonymous memory.

The kernel picks the victim by a number it calls badness (mm/oom_kill.c): the candidate's resident pages, swap entries and page tables, plus its oom_score_adj (-1000 to 1000) counted in thousandths of the memory the OOM is about (the cgroup limit here, RAM plus swap for a global OOM). The highest badness dies; -1000 exempts a task completely.

/proc/PID/oom_score, which choom printed above, shows the same idea on a machine-wide scale: (1000 + the task's share of RAM plus swap in thousandths + its adjustment) × 2/3. The swapping hog held about 137 MiB (RAM, swap, program files and page tables) of about 4.3 GiB of RAM plus swap, 31 thousandths, and 1031 × 2/3 rounds down to 687. An idle, unadjusted task scores 666.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --unit=k-mem-cont --uid=$USER -p MemoryMax=64M -p MemorySwapMax=0 -p OOMPolicy=continue \ sh -c '/var/tmp/k-mem-hog 128; exec sleep 600'
Running as unit: k-mem-cont.service; invocation ID: e2d4f3d1462d48d5921bb06fcdb4a0dc
$ systemctl is-active k-mem-cont grep . /sys/fs/cgroup/system.slice/k-mem-cont.service/memory.events
active low 0 high 0 max 54 oom 1 oom_kill 1 oom_group_kill 0 sock_throttled 0
$ sudo systemctl stop k-mem-cont

With OOMPolicy=continue the unit stays active after the kill, and memory.events holds the evidence: max counts the times usage hit the limit, oom the OOM events and oom_kill the processes killed. continue suits a service whose main process replaces its own workers; OOMPolicy=kill instead sets memory.oom.group, so the kernel kills the whole cgroup together.

deploy@web01 · Ubuntu 26.04 LTS
$ systemctl show -p OOMScoreAdjust ssh systemd-journald systemd-udevd
OOMScoreAdjust=0 OOMScoreAdjust=-250 OOMScoreAdjust=-1000

OOMScoreAdjust= sets oom_score_adj for a unit's processes. systemd's own units protect the journal (-250) and exempt udev (-1000); ssh.service is not adjusted. Protect a process only if its death takes the machine down with it: every protected process makes another one the victim.

deploy@web01 · Ubuntu 26.04 LTS
$ systemctl show -p DefaultOOMPolicy systemctl status systemd-oomd
DefaultOOMPolicy=stop Unit systemd-oomd.service could not be found.

systemd-oomd is a userspace daemon that kills whole cgroups when memory pressure or swap use stays high, before the kernel would act. Ubuntu Desktop and Fedora enable it; Ubuntu Server 26.04 and RHEL 10 do not install it. On servers the kernel OOM killer, cgroup limits and OOMPolicy= decide, so memory.events and journalctl -k are where you look.

The RHEL 10 kernel has the same memory defaults, and no /proc/pressure unless it boots with psi=1, as the method lesson showed. These are the kernel's own values because this Rocky cloud image has no TuneD. A standard RHEL 10 install applies a TuneD profile (virtual-guest on a VM, throughput-performance on most servers) that changes vm.swappiness and the dirty ratios, so run tuned-adm active before reading them as defaults; the tuning lesson shows what a profile changes.

deploy@rocky10 · Rocky Linux 10.2
$ sysctl vm.overcommit_memory vm.swappiness vm.dirty_background_ratio vm.dirty_ratio
vm.overcommit_memory = 0 vm.swappiness = 60 vm.dirty_background_ratio = 10 vm.dirty_ratio = 20
$ rpm -q tuned
package tuned is not installed
$ systemctl show -p DefaultOOMPolicy systemctl status systemd-oomd
DefaultOOMPolicy=stop Unit systemd-oomd.service could not be found.
$ cat /proc/pressure/memory
cat: /proc/pressure/memory: No such file or directory

Try this

Predict the victim before you look. Run a 200 MiB hog and then a 150 MiB one with the maximum adjustment in a 300 MiB unit without swap: sudo systemd-run --unit=k-mem-pick --uid=$USER -p MemoryMax=300M -p MemorySwapMax=0 -p OOMPolicy=continue sh -c '/var/tmp/k-mem-hog 200 & sleep 2; choom -n 1000 -- /var/tmp/k-mem-hog 150; exec sleep 600'. Which one does the kernel kill? Expect journalctl -u k-mem-pick -I -o cat to show holding 200 MiB, a Killed from sh and systemd's OOM note, the Killed process line in journalctl -k to end with oom_score_adj:1000, and pgrep -a k-mem-hog to still find the 200 MiB hog: the adjustment added the whole 300 MiB limit to the smaller one's badness. Finish with sudo systemctl stop k-mem-pick and delete the three scripts and /var/tmp/k-mem.dat.

Takeaway

Size services by their cgroup, never by summed RSS: measure memory.peak over a representative period, add headroom, and set MemoryHigh= there first so that memory.events shows pressure before anything dies. Add MemoryMax= above it as the last line of defence only together with the OOMPolicy= and Restart= you want when it fires, and prove what happened with memory.events and journalctl -k.

Quick check
01A dashboard adds up the RSS of 40 php-fpm workers and reports 6 GB on a server with 4 GB of RAM that is not swapping. What explains the number?
Incorrect — Nothing is compressed here, and memory that does not exist cannot be resident. The sum itself is wrong.
Incorrect — That describes VSZ. RSS counts only pages that are resident in RAM.
Correct — Copy-on-write pages and shared libraries count in every process that maps them. PSS or the cgroup's memory.current gives the real total.
Incorrect — RSS is read from the kernel's counters at the moment you ask; it does not lag.
02On a host with 8 GB of swap, a batch service has MemoryMax=2G. It gets slower every night but is never OOM-killed, and its cgroup's memory.swap.current reads 3 GB. What is happening?
Correct — memory.max limits resident memory, and memory.swap.max defaults to unlimited, so reclaim swaps the service instead of killing it.
Incorrect — A cgroup limit covers every process in the cgroup, children included.
Incorrect — memory.max is a hard limit on the cgroup whatever the rest of the machine is doing.
Incorrect — The default is stop, and no OOMPolicy value disables the kernel's OOM killer.
03A service with the default OOMPolicy runs a supervisor process and four workers. The kernel OOM-kills one worker inside the unit's MemoryMax. What happens next?
Incorrect — That is what OOMPolicy=continue allows. The default policy does not leave the unit running.
Incorrect — systemd watches the cgroup's OOM events, not only the main process, and reacts to any kill in the unit.
Incorrect — memory.oom.group is set only by OOMPolicy=kill; the kernel kills one task by default.
Correct — DefaultOOMPolicy=stop stops the unit after any kernel OOM kill in it; Restart= then decides whether it comes back.

Related