Namespaces and cgroups by hand

unshare, nsenter and raw cgroup files.

Advanced16 min · lesson 10 of 21

A process in a container is OOM-killed or cannot start a thread while the host has gigabytes free and ps on the host shows nothing unusual. To explain that you have to read the container the way the kernel sees it. A container runtime builds a container from two kernel features. Namespaces give a process its own instance of one kind of global resource, so it sees its own hostname, process list or network stack. Control groups (cgroups) account for and limit what a group of processes may consume, and their counters record every time a limit fired. In this lesson you create each namespace type with unshare, see how a user namespace maps user IDs and why Ubuntu stops an ordinary user from mapping itself to root, inspect a running namespaced process from the host, and then limit it by writing cgroup v2 files yourself until the memory and task limits bite and their counters show it. The previous lesson set limits through systemd; here you use the kernel interface that systemd and container runtimes write to underneath. The security view of containers from the host is in the Advanced Linux security course (optional, not a prerequisite for this one), in its lesson "Containers and namespaces from the host".

Eight namespace types

Every process belongs to exactly one namespace of each type, and /proc/PID/ns has one symbolic link per type. The number in brackets is the inode number of the namespace on the kernel's internal nsfs filesystem: it identifies the namespace, so two processes that show the same number for a type share that namespace.

deploy@web01 · Ubuntu 26.04 LTS
$ ls -l /proc/$$/ns
total 0 lrwxrwxrwx 1 deploy deploy 0 Sep 27 09:25 cgroup -> cgroup:[4026531835] lrwxrwxrwx 1 deploy deploy 0 Sep 27 09:25 ipc -> ipc:[4026531839] lrwxrwxrwx 1 deploy deploy 0 Sep 27 09:25 mnt -> mnt:[4026531832] lrwxrwxrwx 1 deploy deploy 0 Sep 27 09:25 net -> net:[4026531833] lrwxrwxrwx 1 deploy deploy 0 Sep 27 09:25 pid -> pid:[4026531836] lrwxrwxrwx 1 deploy deploy 0 Sep 27 09:25 pid_for_children -> pid:[4026531836] lrwxrwxrwx 1 deploy deploy 0 Sep 27 09:25 time -> time:[4026531834] lrwxrwxrwx 1 deploy deploy 0 Sep 27 09:25 time_for_children -> time:[4026531834] lrwxrwxrwx 1 deploy deploy 0 Sep 27 09:25 user -> user:[4026531837] lrwxrwxrwx 1 deploy deploy 0 Sep 27 09:25 uts -> uts:[4026531838]

Eight types: cgroup (which part of the cgroup tree the process sees as its root), ipc (System V IPC objects and POSIX message queues), mnt (the mount table), net (interfaces, routes, firewall rules, sockets), pid (process IDs), time (offsets for the boot and monotonic clocks), user (user and group IDs and the capabilities that go with them) and uts (hostname). pid_for_children and time_for_children name the namespace this process's next child will be created in, which matters below. Most processes share the initial namespaces, the ones PID 1 uses; a service with systemd sandboxing such as PrivateTmp= gets a mount namespace of its own.

A hand-built box and how the host sees it
Namespaces: what it sees
mnt, pid, net
mounts, process IDs, network stack
uts, ipc, time
hostname, System V IPC, boot clocks
user, cgroup
ID mapping and capabilities, cgroup root
cgroup: what it may use
memory.max
the kernel OOM-kills inside the group above it
pids.max
fork() fails with EAGAIN at the limit
cgroup.procs
which processes belong to the group
From the host
/proc/PID/ns/*
one inode number per namespace
nsenter --target PID
run a command inside its namespaces
/proc/PID/root
its files, through its mount namespace

Creating namespaces with unshare

unshare calls the unshare(2) system call for the types you name and then runs a program in the new namespaces. Creating any type except user needs the CAP_SYS_ADMIN capability, so these commands use sudo. Start with the hostname and the network stack.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo unshare --uts sh -c 'hostname k-ns-demo; hostname' hostname
k-ns-demo web01
$ sudo unshare --net ip -brief link
lo DOWN 00:00:00:00:00:00 <LOOPBACK>

The shell in the new UTS namespace changed its hostname to k-ns-demo while the host kept web01. A new network namespace starts with one loopback interface, down, and nothing else: no host interfaces, routes, firewall rules or listening sockets. A service that needs the network there must be given an interface, for example one end of a veth pair.

A PID namespace works differently. The process that calls unshare stays where it is, and only its next child becomes PID 1 in the new namespace, which is why --fork is needed.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo unshare --pid --fork sh -c 'echo "my PID: $$"; ps -e --no-headers | wc -l'
my PID: 1 125
$ sudo unshare --pid --fork --mount-proc ps -ef
UID PID PPID C STIME TTY TIME CMD root 1 0 0 09:25 ? 00:00:00 ps -ef

In the first command the shell is PID 1, but ps still counts all 125 processes on the host: ps reads /proc, and that is still the host's procfs mount, which lists the host's processes. --mount-proc also creates a mount namespace and mounts a fresh /proc inside it, and then ps sees only itself, as PID 1.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo unshare --mount sh -c 'mount -t tmpfs k-ns-scratch /media; findmnt /media' findmnt /media
TARGET SOURCE FSTYPE OPTIONS /media k-ns-scratch tmpfs rw,relatime,inode64
$ ipcmk --queue ipcs -q sudo unshare --ipc ipcs -q
Message queue id: 0 ------ Message Queues -------- key msqid owner perms used-bytes messages 0x337d2c49 0 deploy 644 0 0 ------ Message Queues -------- key msqid owner perms used-bytes messages

The tmpfs mounted on /media exists only in the new mount namespace; the host's findmnt finds nothing and exits 1. unshare makes every mount in the new namespace private, so later mounts do not propagate back to the host. The message queue created with ipcmk is visible on the host and absent in the new IPC namespace. (ipcrm -q with the printed ID removes the queue.) POSIX shared memory under /dev/shm is different: it is a file on a tmpfs mount, so it follows the mount namespace, not the IPC one.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/uptime sudo unshare --time --fork --boottime 864000 cat /proc/uptime date; sudo unshare --time --fork --boottime 864000 date
3564.10 6103.69 867564.10 6103.69 Sun Sep 27 09:25:30 UTC 2026 Sun Sep 27 09:25:30 UTC 2026

The first number in /proc/uptime is seconds since boot. In the new time namespace it is 864,000 seconds (ten days) larger, because --boottime set an offset for the boot clock; --monotonic does the same for the monotonic clock. The wall clock is not namespaced, so date agrees in both. Checkpoint and restore tools use time namespaces to move a container to another host without its clocks jumping. The cgroup namespace appears later, with the box.

User namespaces: IDs and capabilities

A user namespace has its own user and group IDs, mapped to IDs outside it through /proc/PID/uid_map and gid_map. The process that creates one gets a full set of capabilities inside it, but those capabilities count only for namespaces and objects that user namespace owns. It is the one type an ordinary user can create without privileges, and rootless containers are built on it. Try it as deploy:

deploy@web01 · Ubuntu 26.04 LTS
$ unshare --user sh -c "id; cat /proc/self/uid_map; grep CapEff /proc/self/status"
uid=65534(nobody) gid=65534(nogroup) groups=65534(nogroup),65534(nogroup),65534(nogroup) CapEff: 0000000000000000
$ unshare --user --map-root-user id
unshare: write failed /proc/self/uid_map: Operation not permitted
$ sysctl kernel.apparmor_restrict_unprivileged_userns user.max_user_namespaces
kernel.apparmor_restrict_unprivileged_userns = 1 user.max_user_namespaces = 13691

Nothing was written to uid_map, so every ID is unmapped and shows as the overflow ID 65534: nobody, and nogroup three times for deploy's three groups. CapEff is the effective capability set written as a bitmask, one bit per capability, read like the signal masks in the process lifecycle lesson. It is zero because execve() recalculates capabilities, and a process whose user ID in the namespace is not 0 loses them all. --map-root-user would write 0 1001 1 (ID 0 inside is 1001 outside, one ID) to the map, and Ubuntu refuses. The refusal comes from AppArmor, Ubuntu's mandatory access control system, which confines programs according to per-program rule sets called profiles; a program without a profile is "unconfined". With kernel.apparmor_restrict_unprivileged_userns=1, AppArmor moves an unconfined program that creates a user namespace into a profile named unprivileged_userns that denies the capabilities the mapping needs. Programs such as podman that need user namespaces ship their own profile that allows them. RHEL 10 has no such setting and relies on user.max_user_namespaces and SELinux.

deploy@rocky10 · Rocky Linux 10.2
$ sysctl user.max_user_namespaces ls /proc/sys/kernel/apparmor_restrict_unprivileged_userns
user.max_user_namespaces = 14365 ls: cannot access '/proc/sys/kernel/apparmor_restrict_unprivileged_userns': No such file or directory
$ unshare --user --map-root-user sh -c "id; cat /proc/self/uid_map; grep CapEff /proc/self/status"
uid=0(root) gid=0(root) groups=0(root),65534(nobody) context=unconfined_u:unconfined_r:unconfined_t:s0-s0:c0.c1023 0 1001 1 CapEff: 000001ffffffffff
$ unshare --user --map-root-user --net --uts sh -c 'hostname k-ns-user; hostname; ip link set lo up; ip -brief link' hostname
k-ns-user lo UNKNOWN 00:00:00:00:00:00 <LOOPBACK,UP,LOWER_UP> rocky10
$ unshare --user --map-root-user sh -c 'ls -ln /etc/shadow; cat /etc/shadow'
----------. 1 65534 65534 717 Sep 27 09:15 /etc/shadow cat: /etc/shadow: Permission denied

On Rocky Linux the same user maps itself to root: uid_map reads 0 1001 1 and CapEff is full. With those capabilities it set a hostname and brought lo up (UNKNOWN is the normal state of a loopback interface that is up), in UTS and network namespaces that its user namespace owns, while the host kept rocky10. The capabilities stop at the host's objects. /etc/shadow belongs to host UID 0, which is not mapped into this namespace, so it shows as 65534 and the read is refused.

Root can map ranges of IDs, which is how container engines give a container its own block of host IDs.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo unshare --user --map-users=0:200000:65536 --map-groups=0:200000:65536 --setuid 0 --setgid 0 sh -c 'id; cat /proc/self/uid_map; touch /var/tmp/k-ns-userns-file' ls -ln /var/tmp/k-ns-userns-file
uid=0(root) gid=0(root) groups=0(root) 0 200000 65536 -rw-r--r-- 1 200000 200000 0 Sep 27 09:25 /var/tmp/k-ns-userns-file

Inside, the process is uid=0(root); the map says IDs 0 to 65535 are 200000 to 265535 on the host, and the file it created belongs to 200000. Rootless podman and Docker's userns-remap mode take such ranges from /etc/subuid and /etc/subgid, so root in the container is an unprivileged user outside it.

A box, seen from the host

Now combine every type except user into one long-running box. Its PID 1 is a small script that names the box, mounts a private tmpfs on /media and waits. Save it as /var/tmp/k-ns-init and make it executable.

/var/tmp/k-ns-init
#!/bin/sh
# k-ns-init: PID 1 of a hand-built box. Names the box, mounts a private tmpfs, then waits.
trap 'exit 0' TERM
hostname k-ns-box
mount -t tmpfs k-ns-scratch /media
echo 'written inside the box' > /media/marker
while true; do sleep 3600; done

The trap matters. The kernel delivers a signal to the init process of a PID namespace only if it has a handler for it (SIGKILL and SIGSTOP sent from an ancestor namespace are the exceptions), so without the trap the SIGTERM from systemctl stop would be ignored until systemd gave up and sent SIGKILL. Start the box as a transient service; Delegate= and DelegateSubgroup= prepare its cgroup for the next section.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --unit=k-ns-box -p Delegate=yes -p DelegateSubgroup=init unshare --pid --mount --uts --ipc --net --cgroup --time --fork --mount-proc /var/tmp/k-ns-init
Running as unit: k-ns-box.service; invocation ID: 8a0421e1bf7045c681b81069662e4e35
$ systemctl status k-ns-box --lines=0
● k-ns-box.service - [systemd-run] /usr/bin/unshare --pid --mount --uts --ipc --net --cgroup --time --fork --mount-proc /var/tmp/k-ns-init Loaded: loaded (/run/systemd/transient/k-ns-box.service; transient) Transient: yes Active: active (running) since Sun 2026-09-27 09:25:30 UTC; 38ms ago Invocation: 8a0421e1bf7045c681b81069662e4e35 Main PID: 218232 (unshare) Tasks: 3 (limit: 4107) Memory: 2.6M (peak: 3.8M) CPU: 13ms CGroup: /system.slice/k-ns-box.service └─init ├─218232 /usr/bin/unshare --pid --mount --uts --ipc --net --cgroup --time --fork --mount-proc /var/tmp/k-ns-init ├─218243 /bin/sh /var/tmp/k-ns-init └─218246 sleep 3600

unshare is the main process. Its child runs the script, and the kernel named that process k-ns-init after the script file, so pgrep -x k-ns-init finds it.

deploy@web01 · Ubuntu 26.04 LTS
$ grep NSpid /proc/$(pgrep -x k-ns-init)/status
NSpid: 218243 1
$ sudo lsns --task $(pgrep -x k-ns-init)
NS TYPE NPROCS PID USER COMMAND 4026531837 user 125 1 root /usr/lib/systemd/systemd --switched-root --system --deserialize=50 4026532442 mnt 3 218232 root ├─/usr/bin/unshare --pid --mount --uts --ipc --net --cgroup --time --fork --mount-proc /var/tmp/k-ns-init 4026532455 uts 3 218232 root ├─/usr/bin/unshare --pid --mount --uts --ipc --net --cgroup --time --fork --mount-proc /var/tmp/k-ns-init 4026532456 ipc 3 218232 root ├─/usr/bin/unshare --pid --mount --uts --ipc --net --cgroup --time --fork --mount-proc /var/tmp/k-ns-init 4026532457 pid 2 218243 root │ └─/bin/sh /var/tmp/k-ns-init 4026532458 cgroup 3 218232 root ├─/usr/bin/unshare --pid --mount --uts --ipc --net --cgroup --time --fork --mount-proc /var/tmp/k-ns-init 4026532459 net 3 218232 root └─/usr/bin/unshare --pid --mount --uts --ipc --net --cgroup --time --fork --mount-proc /var/tmp/k-ns-init 4026532577 time 2 218243 root └─/bin/sh /var/tmp/k-ns-init

NSpid lists the process's ID in each PID namespace it is visible in, outermost first: host PID 218243 is PID 1 inside. lsns --task shows each of its namespaces with the number of processes in it and the lowest PID. The user namespace is the host's, shared by all 125 processes on the host, so root in this box is the host's root. mnt, uts, ipc, cgroup and net hold three processes, unshare included, because unshare(2) moves the caller into those at once. pid and time hold two, because they apply only to children.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo nsenter --target $(pgrep -x k-ns-init) --all sh -c "hostname; ip -brief link; ps -ef"
k-ns-box lo DOWN 00:00:00:00:00:00 <LOOPBACK> UID PID PPID C STIME TTY TIME CMD root 1 0 0 09:25 ? 00:00:00 /bin/sh /var/tmp/k-ns-init root 4 1 0 09:25 ? 00:00:00 sleep 3600 root 5 0 0 09:25 ? 00:00:00 sh -c hostname; ip -brief link; ps -ef root 8 5 0 09:25 ? 00:00:00 ps -ef

nsenter joins the target's namespaces with setns(2) and runs a command there: the box's hostname, its lone loopback interface and its process list. The sh that nsenter started shows PPID 0, because its parent is outside the namespace and has no PID in it.

deploy@web01 · Ubuntu 26.04 LTS
$ ls /media sudo cat /proc/$(pgrep -x k-ns-init)/root/media/marker
written inside the box
$ sudo grep k-ns-scratch /proc/$(pgrep -x k-ns-init)/mountinfo
413 347 0:60 / /media rw,relatime - tmpfs k-ns-scratch rw,inode64

/media is empty on the host. /proc/PID/root is the process's root directory as seen through its own mount namespace, so the host can read a file that exists only inside the box without entering it, and /proc/PID/mountinfo lists the box's mount table. That is how you collect files and mount details from a container whose image has no shell.

cgroup v2 by hand

All cgroups live in one tree mounted at /sys/fs/cgroup, and on a systemd host systemd is its single writer: it creates a cgroup per unit and expects nobody else to reorganise them. Delegate=yes hands a unit's subtree to the program running in it, which is how container managers and the per-user systemd instance get a subtree of their own. (mkdir directly under /sys/fs/cgroup also works, but it bypasses systemd; keep that for a throwaway machine.)

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/$(pgrep -x k-ns-init)/cgroup sudo nsenter --target $(pgrep -x k-ns-init) --all --join-cgroup cat /proc/self/cgroup
0::/system.slice/k-ns-box.service/init 0::/

From the host the box's cgroup is /system.slice/k-ns-box.service/init. Inside, --join-cgroup put the command in that same cgroup and the cgroup namespace makes it look like the root, /.

deploy@web01 · Ubuntu 26.04 LTS
$ cd /sys/fs/cgroup/system.slice/k-ns-box.service cat cgroup.controllers cgroup.subtree_control cat init/cgroup.procs
cpuset cpu io memory pids 218232 218243 218246

cgroup.controllers lists the controllers this cgroup may use; cgroup.subtree_control, empty here, lists those it passes on to its children. Two rules follow. A child has files for a controller, such as memory.max, only if its parent has enabled that controller in subtree_control. And a cgroup that contains processes cannot enable controllers for children (the root cgroup excepted), which is why DelegateSubgroup=init started the unit's processes in a child named init and left the unit's own cgroup empty.

deploy@web01 · Ubuntu 26.04 LTS
$ cd /sys/fs/cgroup/system.slice/k-ns-box.service echo '+memory +pids' | sudo tee cgroup.subtree_control sudo mkdir box echo 64M | sudo tee box/memory.max echo 0 | sudo tee box/memory.swap.max echo 8 | sudo tee box/pids.max cat box/memory.max
+memory +pids 64M 0 8 67108864
$ echo '+memory +pids' | sudo tee /sys/fs/cgroup/system.slice/k-ns-box.service/init/cgroup.subtree_control
+memory +pids /sys/fs/cgroup/system.slice/k-ns-box.service/init/cgroup.subtree_control: Device or resource busy (os error 16)

tee echoes what it writes. memory.max accepts 64M and stores bytes, so read the value back rather than trusting what you wrote. memory.swap.max is 0 so that the group cannot use swap either; this machine has none, but on a host with swap the process would be paged out before anything was killed. The same write into init failed with Device or resource busy (EBUSY, which Ubuntu's Rust tee prints as os error 16): the second rule. Now move PID 1 of the box into the limited cgroup by writing its PID to cgroup.procs.

deploy@web01 · Ubuntu 26.04 LTS
$ cd /sys/fs/cgroup/system.slice/k-ns-box.service echo $(pgrep -x k-ns-init) | sudo tee box/cgroup.procs grep . init/cgroup.procs box/cgroup.procs
218243 init/cgroup.procs:218232 init/cgroup.procs:218246 box/cgroup.procs:218243

Only the process moved. Its existing child, the sleep, stayed in init; children it creates from now on start in box. To test the task limit, this program forks children that sleep for five seconds until fork() fails. Save it as /var/tmp/k-ns-fork, make it executable and run it inside the box and its cgroup.

/var/tmp/k-ns-fork
#!/usr/bin/python3
# k-ns-fork: starts children that sleep for 5 s until fork() fails (at most 20), then reaps them.
import os, time
children = []
try:
while len(children) < 20:
pid = os.fork()
if pid == 0:
time.sleep(5)
os._exit(0)
children.append(pid)
except OSError as err:
print(len(children), 'children started, then:', err)
for pid in children:
os.waitpid(pid, 0)
deploy@web01 · Ubuntu 26.04 LTS
$ sudo nsenter --target $(pgrep -x k-ns-init) --all --join-cgroup /var/tmp/k-ns-fork
5 children started, then: [Errno 11] Resource temporarily unavailable
$ cd /sys/fs/cgroup/system.slice/k-ns-box.service/box grep . pids.current pids.peak pids.max pids.events
pids.current:1 pids.peak:8 pids.max:8 pids.events:max 1

pids.max is 8. k-ns-init, nsenter and the Python process took three, five children took the rest, and the sixth fork() failed with EAGAIN, errno 11. pids.peak records the high-water mark, and max in pids.events counts the times the group hit its limit. Now the memory limit.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo nsenter --target $(pgrep -x k-ns-init) --all --join-cgroup python3 -c "data = b'x' * (100 << 20)"
Killed
$ cat /sys/fs/cgroup/system.slice/k-ns-box.service/box/memory.events
low 0 high 0 max 52 oom 1 oom_kill 1 oom_group_kill 0 sock_throttled 0
$ journalctl -k -n 30 -o cat | grep -E "Memory cgroup out of memory|oom-kill:"
oom-kill:constraint=CONSTRAINT_MEMCG,nodemask=(null),cpuset=k-ns-box.service,mems_allowed=0,oom_memcg=/system.slice/k-ns-box.service/box,task_memcg=/system.slice/k-ns-box.service/box,task=python3,pid=218645,uid=0 Memory cgroup out of memory: Killed process 218645 (python3) total-vm:122400kB, anon-rss:65264kB, file-rss:5696kB, shmem-rss:0kB, UID:0 pgtables:200kB oom_score_adj:0

Python asked for 100 MiB in a group limited to 64 MiB. The kernel reclaimed what it could (max counts the times usage hit the limit), found nothing more to free, and the OOM killer ended the process in the group with the highest OOM score, Python, recorded in oom and oom_kill. The kernel log line CONSTRAINT_MEMCG with oom_memcg=/system.slice/k-ns-box.service/box says this was a cgroup OOM, not a machine out of memory; the memory lesson reads that report in full. That is the symptom from the start of this lesson explained: the host had memory to spare, and the limit that fired belongs to the box.

deploy@web01 · Ubuntu 26.04 LTS
$ systemctl is-active k-ns-box systemctl show -p Delegate -p OOMPolicy k-ns-box
active OOMPolicy=continue Delegate=yes

The box is still running, and not because the kill was small. systemd saw the OOM kill in the unit's subtree, since memory.events counts upwards through the tree, and the memory lesson showed that the default policy then stops the whole unit. But a unit with Delegate=yes gets OOMPolicy=continue by default (systemd.service(5)): the program it delegated to is expected to deal with kills in its own subtree, as a container manager does. Without delegation, the same kill would have stopped the box.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop k-ns-box pgrep -x k-ns-init || echo "no k-ns-init left" ls /sys/fs/cgroup/system.slice/k-ns-box.service
no k-ns-init left ls: cannot access '/sys/fs/cgroup/system.slice/k-ns-box.service': No such file or directory

Stopping the unit killed every process in its subtree, and systemd removed its cgroup directory, including the box group you made by hand. The namespaces went with their last process. A namespace outlives its processes only while something holds it open: a bind mount of its /proc/PID/ns file (ip netns add and unshare --net=FILE do this) or an open file descriptor.

Try this

Start the box again with the same systemd-run command, create the box group and move k-ns-init into it as above. Then try sudo rmdir /sys/fs/cgroup/system.slice/k-ns-box.service/box: it fails with Device or resource busy, because a cgroup with a process in it cannot be removed. Move k-ns-init back by writing its PID to init/cgroup.procs, check that box/pids.current reads 0, and run the rmdir again; it succeeds. Finish with sudo systemctl stop k-ns-box and remove /var/tmp/k-ns-init and /var/tmp/k-ns-fork.

Takeaway

To understand a container, read it from the host: its namespaces in /proc/PID/ns and lsns, its files through /proc/PID/root, and its limits and counters in its cgroup directory, where pids.events and memory.events tell you whether a limit actually fired.

Quick check
01On an Ubuntu 26.04 server, a developer runs unshare --user --map-root-user id as an ordinary user and gets "write failed /proc/self/uid_map: Operation not permitted". The same command works on their RHEL 10 laptop VM. What explains the difference?
Incorrect — On Ubuntu the limit is in the thousands, and unshare --user id without the mapping succeeds, so namespace creation is not what fails.
Incorrect — An ordinary user may map its own UID into a user namespace it created; the RHEL VM shows that working without sudo.
Correct — With kernel.apparmor_restrict_unprivileged_userns=1, an unconfined program's new user namespace gets the unprivileged_userns profile, and the uid_map write that needs a capability is refused.
Incorrect — Ubuntu's kernel supports user namespaces; root and profiled programs such as podman use them.
02You started a box with sudo unshare --pid --fork --net ... /var/tmp/k-ns-init. lsns --task shows its net namespace with 3 processes but its pid namespace with only 2. Why the difference?
Correct — unshare(2) moves the caller into most new namespaces at once; the PID (and time) namespace is used for the children the caller creates after the call.
Incorrect — Kernel threads live in the initial namespaces and lsns does not count them in a new network namespace.
Incorrect — A process's PID namespace is fixed when it is created; nothing moves it later.
Incorrect — No nsenter was running; the extra process is unshare, which the tree output names.
03You wrote a PID into box/cgroup.procs and set box/pids.max to 8, but a long-running child that the process started earlier keeps running outside the limit. What happened?
Incorrect — The limit applies to every task in the cgroup at any time; the question is which tasks are in the cgroup.
Correct — Migration moves the named process (with all its threads). Children created before the move keep their cgroup, and only new children inherit the new one.
Incorrect — cgroup membership is independent of namespaces; the child was simply never moved.
Incorrect — The move is immediate, as cgroup.procs reading back the PID shows; subtree_control governs controllers, not membership.

Related