Namespaces and cgroups by hand
unshare, nsenter and raw cgroup files.
A process in a container is OOM-killed or cannot start a thread while the host has gigabytes free and ps on the host shows nothing unusual. To explain that you have to read the container the way the kernel sees it. A container runtime builds a container from two kernel features. Namespaces give a process its own instance of one kind of global resource, so it sees its own hostname, process list or network stack. Control groups (cgroups) account for and limit what a group of processes may consume, and their counters record every time a limit fired. In this lesson you create each namespace type with unshare, see how a user namespace maps user IDs and why Ubuntu stops an ordinary user from mapping itself to root, inspect a running namespaced process from the host, and then limit it by writing cgroup v2 files yourself until the memory and task limits bite and their counters show it. The previous lesson set limits through systemd; here you use the kernel interface that systemd and container runtimes write to underneath. The security view of containers from the host is in the Advanced Linux security course (optional, not a prerequisite for this one), in its lesson "Containers and namespaces from the host".
Eight namespace types
Every process belongs to exactly one namespace of each type, and /proc/PID/ns has one symbolic link per type. The number in brackets is the inode number of the namespace on the kernel's internal nsfs filesystem: it identifies the namespace, so two processes that show the same number for a type share that namespace.
Eight types: cgroup (which part of the cgroup tree the process sees as its root), ipc (System V IPC objects and POSIX message queues), mnt (the mount table), net (interfaces, routes, firewall rules, sockets), pid (process IDs), time (offsets for the boot and monotonic clocks), user (user and group IDs and the capabilities that go with them) and uts (hostname). pid_for_children and time_for_children name the namespace this process's next child will be created in, which matters below. Most processes share the initial namespaces, the ones PID 1 uses; a service with systemd sandboxing such as PrivateTmp= gets a mount namespace of its own.
Creating namespaces with unshare
unshare calls the unshare(2) system call for the types you name and then runs a program in the new namespaces. Creating any type except user needs the CAP_SYS_ADMIN capability, so these commands use sudo. Start with the hostname and the network stack.
The shell in the new UTS namespace changed its hostname to k-ns-demo while the host kept web01. A new network namespace starts with one loopback interface, down, and nothing else: no host interfaces, routes, firewall rules or listening sockets. A service that needs the network there must be given an interface, for example one end of a veth pair.
A PID namespace works differently. The process that calls unshare stays where it is, and only its next child becomes PID 1 in the new namespace, which is why --fork is needed.
In the first command the shell is PID 1, but ps still counts all 125 processes on the host: ps reads /proc, and that is still the host's procfs mount, which lists the host's processes. --mount-proc also creates a mount namespace and mounts a fresh /proc inside it, and then ps sees only itself, as PID 1.
The tmpfs mounted on /media exists only in the new mount namespace; the host's findmnt finds nothing and exits 1. unshare makes every mount in the new namespace private, so later mounts do not propagate back to the host. The message queue created with ipcmk is visible on the host and absent in the new IPC namespace. (ipcrm -q with the printed ID removes the queue.) POSIX shared memory under /dev/shm is different: it is a file on a tmpfs mount, so it follows the mount namespace, not the IPC one.
The first number in /proc/uptime is seconds since boot. In the new time namespace it is 864,000 seconds (ten days) larger, because --boottime set an offset for the boot clock; --monotonic does the same for the monotonic clock. The wall clock is not namespaced, so date agrees in both. Checkpoint and restore tools use time namespaces to move a container to another host without its clocks jumping. The cgroup namespace appears later, with the box.
User namespaces: IDs and capabilities
A user namespace has its own user and group IDs, mapped to IDs outside it through /proc/PID/uid_map and gid_map. The process that creates one gets a full set of capabilities inside it, but those capabilities count only for namespaces and objects that user namespace owns. It is the one type an ordinary user can create without privileges, and rootless containers are built on it. Try it as deploy:
Nothing was written to uid_map, so every ID is unmapped and shows as the overflow ID 65534: nobody, and nogroup three times for deploy's three groups. CapEff is the effective capability set written as a bitmask, one bit per capability, read like the signal masks in the process lifecycle lesson. It is zero because execve() recalculates capabilities, and a process whose user ID in the namespace is not 0 loses them all. --map-root-user would write 0 1001 1 (ID 0 inside is 1001 outside, one ID) to the map, and Ubuntu refuses. The refusal comes from AppArmor, Ubuntu's mandatory access control system, which confines programs according to per-program rule sets called profiles; a program without a profile is "unconfined". With kernel.apparmor_restrict_unprivileged_userns=1, AppArmor moves an unconfined program that creates a user namespace into a profile named unprivileged_userns that denies the capabilities the mapping needs. Programs such as podman that need user namespaces ship their own profile that allows them. RHEL 10 has no such setting and relies on user.max_user_namespaces and SELinux.
On Rocky Linux the same user maps itself to root: uid_map reads 0 1001 1 and CapEff is full. With those capabilities it set a hostname and brought lo up (UNKNOWN is the normal state of a loopback interface that is up), in UTS and network namespaces that its user namespace owns, while the host kept rocky10. The capabilities stop at the host's objects. /etc/shadow belongs to host UID 0, which is not mapped into this namespace, so it shows as 65534 and the read is refused.
Root can map ranges of IDs, which is how container engines give a container its own block of host IDs.
Inside, the process is uid=0(root); the map says IDs 0 to 65535 are 200000 to 265535 on the host, and the file it created belongs to 200000. Rootless podman and Docker's userns-remap mode take such ranges from /etc/subuid and /etc/subgid, so root in the container is an unprivileged user outside it.
A box, seen from the host
Now combine every type except user into one long-running box. Its PID 1 is a small script that names the box, mounts a private tmpfs on /media and waits. Save it as /var/tmp/k-ns-init and make it executable.
#!/bin/sh# k-ns-init: PID 1 of a hand-built box. Names the box, mounts a private tmpfs, then waits.trap 'exit 0' TERMhostname k-ns-boxmount -t tmpfs k-ns-scratch /mediaecho 'written inside the box' > /media/markerwhile true; do sleep 3600; done
The trap matters. The kernel delivers a signal to the init process of a PID namespace only if it has a handler for it (SIGKILL and SIGSTOP sent from an ancestor namespace are the exceptions), so without the trap the SIGTERM from systemctl stop would be ignored until systemd gave up and sent SIGKILL. Start the box as a transient service; Delegate= and DelegateSubgroup= prepare its cgroup for the next section.
unshare is the main process. Its child runs the script, and the kernel named that process k-ns-init after the script file, so pgrep -x k-ns-init finds it.
NSpid lists the process's ID in each PID namespace it is visible in, outermost first: host PID 218243 is PID 1 inside. lsns --task shows each of its namespaces with the number of processes in it and the lowest PID. The user namespace is the host's, shared by all 125 processes on the host, so root in this box is the host's root. mnt, uts, ipc, cgroup and net hold three processes, unshare included, because unshare(2) moves the caller into those at once. pid and time hold two, because they apply only to children.
nsenter joins the target's namespaces with setns(2) and runs a command there: the box's hostname, its lone loopback interface and its process list. The sh that nsenter started shows PPID 0, because its parent is outside the namespace and has no PID in it.
/media is empty on the host. /proc/PID/root is the process's root directory as seen through its own mount namespace, so the host can read a file that exists only inside the box without entering it, and /proc/PID/mountinfo lists the box's mount table. That is how you collect files and mount details from a container whose image has no shell.
cgroup v2 by hand
All cgroups live in one tree mounted at /sys/fs/cgroup, and on a systemd host systemd is its single writer: it creates a cgroup per unit and expects nobody else to reorganise them. Delegate=yes hands a unit's subtree to the program running in it, which is how container managers and the per-user systemd instance get a subtree of their own. (mkdir directly under /sys/fs/cgroup also works, but it bypasses systemd; keep that for a throwaway machine.)
From the host the box's cgroup is /system.slice/k-ns-box.service/init. Inside, --join-cgroup put the command in that same cgroup and the cgroup namespace makes it look like the root, /.
cgroup.controllers lists the controllers this cgroup may use; cgroup.subtree_control, empty here, lists those it passes on to its children. Two rules follow. A child has files for a controller, such as memory.max, only if its parent has enabled that controller in subtree_control. And a cgroup that contains processes cannot enable controllers for children (the root cgroup excepted), which is why DelegateSubgroup=init started the unit's processes in a child named init and left the unit's own cgroup empty.
tee echoes what it writes. memory.max accepts 64M and stores bytes, so read the value back rather than trusting what you wrote. memory.swap.max is 0 so that the group cannot use swap either; this machine has none, but on a host with swap the process would be paged out before anything was killed. The same write into init failed with Device or resource busy (EBUSY, which Ubuntu's Rust tee prints as os error 16): the second rule. Now move PID 1 of the box into the limited cgroup by writing its PID to cgroup.procs.
Only the process moved. Its existing child, the sleep, stayed in init; children it creates from now on start in box. To test the task limit, this program forks children that sleep for five seconds until fork() fails. Save it as /var/tmp/k-ns-fork, make it executable and run it inside the box and its cgroup.
#!/usr/bin/python3# k-ns-fork: starts children that sleep for 5 s until fork() fails (at most 20), then reaps them.import os, timechildren = []try:while len(children) < 20:pid = os.fork()if pid == 0:time.sleep(5)os._exit(0)children.append(pid)except OSError as err:print(len(children), 'children started, then:', err)for pid in children:os.waitpid(pid, 0)
pids.max is 8. k-ns-init, nsenter and the Python process took three, five children took the rest, and the sixth fork() failed with EAGAIN, errno 11. pids.peak records the high-water mark, and max in pids.events counts the times the group hit its limit. Now the memory limit.
Python asked for 100 MiB in a group limited to 64 MiB. The kernel reclaimed what it could (max counts the times usage hit the limit), found nothing more to free, and the OOM killer ended the process in the group with the highest OOM score, Python, recorded in oom and oom_kill. The kernel log line CONSTRAINT_MEMCG with oom_memcg=/system.slice/k-ns-box.service/box says this was a cgroup OOM, not a machine out of memory; the memory lesson reads that report in full. That is the symptom from the start of this lesson explained: the host had memory to spare, and the limit that fired belongs to the box.
The box is still running, and not because the kill was small. systemd saw the OOM kill in the unit's subtree, since memory.events counts upwards through the tree, and the memory lesson showed that the default policy then stops the whole unit. But a unit with Delegate=yes gets OOMPolicy=continue by default (systemd.service(5)): the program it delegated to is expected to deal with kills in its own subtree, as a container manager does. Without delegation, the same kill would have stopped the box.
Stopping the unit killed every process in its subtree, and systemd removed its cgroup directory, including the box group you made by hand. The namespaces went with their last process. A namespace outlives its processes only while something holds it open: a bind mount of its /proc/PID/ns file (ip netns add and unshare --net=FILE do this) or an open file descriptor.
Try this
Start the box again with the same systemd-run command, create the box group and move k-ns-init into it as above. Then try sudo rmdir /sys/fs/cgroup/system.slice/k-ns-box.service/box: it fails with Device or resource busy, because a cgroup with a process in it cannot be removed. Move k-ns-init back by writing its PID to init/cgroup.procs, check that box/pids.current reads 0, and run the rmdir again; it succeeds. Finish with sudo systemctl stop k-ns-box and remove /var/tmp/k-ns-init and /var/tmp/k-ns-fork.
Takeaway
To understand a container, read it from the host: its namespaces in /proc/PID/ns and lsns, its files through /proc/PID/root, and its limits and counters in its cgroup directory, where pids.events and memory.events tell you whether a limit actually fired.
unshare --user --map-root-user id as an ordinary user and gets "write failed /proc/self/uid_map: Operation not permitted". The same command works on their RHEL 10 laptop VM. What explains the difference?unshare --user id without the mapping succeeds, so namespace creation is not what fails.sudo unshare --pid --fork --net ... /var/tmp/k-ns-init. lsns --task shows its net namespace with 3 processes but its pid namespace with only 2. Why the difference?