Namespaces from first principles

What each namespace isolates and how to see it with lsns and nsenter.

Advanced15 min · lesson 1 of 24

Run sudo lsns against the main process of an ordinary nginx container and you get eight lines. Seven name a namespace that exists only for that container. The eighth, the user namespace, carries the same number as the one systemd runs in. Most of what this course does starts from that one shared line, and from the fact that the kernel has no container object at all: there are processes, and there are namespaces, cgroups, capabilities and security profiles attached to them.

This course, Advanced container security, assumes Docker for beginners and the image, BuildKit, engine, storage, networking and production/CI parts of Docker in depth, in particular "Engine architecture on Docker 29" and "Networks, drivers and DNS". It does not assume Swarm or the recipes. Commands run on the main lab VM, secopslog-docker, built from the lab kit as "Installing Docker Engine and running your first container" in Docker for beginners describes (./setup/create-lab.sh, then multipass shell secopslog-docker). Lessons that change the daemon or the host, such as rootless Docker, user-namespace remapping, AppArmor profiles and the escape demonstrations, use the second VM, secopslog-docker-sec (./setup/create-lab.sh --profile sec; rebuild it with --recreate when you want it clean), and say so in their first lines. Host-side inspection uses sudo inside the VM, where the ubuntu account has it without a password. Each lesson that has files shows a Lesson files box with the command that fetches and unpacks them inside the VM, into ~/lab/<lesson>. This lesson needs none and changes nothing on the host.

Eight namespace types

A namespace wraps one kind of kernel resource so that the processes inside it get their own instance: their own process ID numbers, their own mount table, their own network stack. Every process belongs to exactly one namespace of each type, and /proc/<pid>/ns/ has a link per type. Your login shell on the VM is in the host's initial set:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ ls -l /proc/self/ns
total 0 lrwxrwxrwx 1 ubuntu ubuntu 0 Oct 8 00:38 cgroup -> 'cgroup:[4026531835]' lrwxrwxrwx 1 ubuntu ubuntu 0 Oct 8 00:38 ipc -> 'ipc:[4026531839]' lrwxrwxrwx 1 ubuntu ubuntu 0 Oct 8 00:38 mnt -> 'mnt:[4026531832]' lrwxrwxrwx 1 ubuntu ubuntu 0 Oct 8 00:38 net -> 'net:[4026531833]' lrwxrwxrwx 1 ubuntu ubuntu 0 Oct 8 00:38 pid -> 'pid:[4026531836]' lrwxrwxrwx 1 ubuntu ubuntu 0 Oct 8 00:38 pid_for_children -> 'pid:[4026531836]' lrwxrwxrwx 1 ubuntu ubuntu 0 Oct 8 00:38 time -> 'time:[4026531834]' lrwxrwxrwx 1 ubuntu ubuntu 0 Oct 8 00:38 time_for_children -> 'time:[4026531834]' lrwxrwxrwx 1 ubuntu ubuntu 0 Oct 8 00:38 user -> 'user:[4026531837]' lrwxrwxrwx 1 ubuntu ubuntu 0 Oct 8 00:38 uts -> 'uts:[4026531838]'

The number in brackets is the inode number of the namespace on the kernel's internal nsfs filesystem. Same number, same namespace; a different number is a different namespace. That comparison is the whole detection method in this lesson. The initial namespaces have fixed numbers in the kernel source, so pid:[4026531836] reads the same on every Linux host. On recent kernels, including this VM's 7.0, that holds for all eight, mount (4026531832) and network (4026531833) included; older kernels allocate those two at boot, so they vary there. pid_for_children and time_for_children exist because those two types take effect for children: a process that calls unshare for them moves its future children, not itself.

What a default container gets

Start one nginx container and look at it from the host. docker inspect gives you its main process ID on the host, which is what every host-side tool wants:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-ns nginx:1.30-alpine
f87f04b646cd46cd419a639389075b2cc9ee122105dd4937912aad8c89516e08
$ P=$(docker inspect -f '{{.State.Pid}}' lab-ns) sudo lsns -p "$P"
NS TYPE NPROCS PID USER COMMAND 4026531837 user 173 1 root /usr/lib/systemd/systemd --switched-root -- 4026532325 mnt 5 324701 root nginx: master process nginx -g daemon off; 4026532326 uts 5 324701 root nginx: master process nginx -g daemon off; 4026532327 ipc 5 324701 root nginx: master process nginx -g daemon off; 4026532328 pid 5 324701 root nginx: master process nginx -g daemon off; 4026532330 cgroup 5 324701 root nginx: master process nginx -g daemon off; 4026532331 net 5 324701 root nginx: master process nginx -g daemon off; 4026532437 time 5 324701 root nginx: master process nginx -g daemon off;

lsns -p lists the namespaces of one process. Seven rows belong to the nginx master process (five processes share them: the master and four workers). The user row lists systemd as the first process in it, because the container never got a user namespace of its own. Your process IDs and namespace numbers will differ; the pattern will not. The same comparison as a loop over all eight types, host PID 1 against the container:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ P=$(docker inspect -f '{{.State.Pid}}' lab-ns) for ns in cgroup ipc mnt net pid time user uts; do printf '%-7s host %-22s container %s\n' "$ns" "$(sudo readlink /proc/1/ns/$ns)" "$(sudo readlink /proc/$P/ns/$ns)" done
cgroup host cgroup:[4026531835] container cgroup:[4026532330] ipc host ipc:[4026531839] container ipc:[4026532327] mnt host mnt:[4026531832] container mnt:[4026532325] net host net:[4026531833] container net:[4026532331] pid host pid:[4026531836] container pid:[4026532328] time host time:[4026531834] container time:[4026532437] user host user:[4026531837] container user:[4026531837] uts host uts:[4026531838] container uts:[4026532326]

Seven numbers differ and user matches. Docker 29.8.2 with runc 1.5.1 also gives each container its own time namespace; it does not change the clocks (both offsets are zero), so you will rarely notice it:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ P=$(docker inspect -f '{{.State.Pid}}' lab-ns) sudo cat /proc/$P/timens_offsets
monotonic 0 0 boottime 0 0

Docker records the choices in the container's HostConfig. Empty PidMode and UTSMode mean private; ipc=private is the daemon's default IPC mode; cgroupns=private is the default on cgroup v2 hosts; an empty UsernsMode means the host user namespace. These fields are what an audit script reads, since they cannot be faked from inside the container:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker inspect -f 'pid={{.HostConfig.PidMode}} ipc={{.HostConfig.IpcMode}} net={{.HostConfig.NetworkMode}} uts={{.HostConfig.UTSMode}} cgroupns={{.HostConfig.CgroupnsMode}} userns={{.HostConfig.UsernsMode}}' lab-ns
pid= ipc=private net=bridge uts= cgroupns=private userns=

From the inside, the private namespaces show up as a short hostname (the container ID), a cgroup path of 0::/ (the cgroup namespace makes the container's own cgroup look like the root of the tree) and a process list that starts with nginx as PID 1:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker exec lab-ns sh -c 'hostname; cat /proc/self/cgroup; ps -o pid,user,comm'
f87f04b646cd 0::/ PID USER COMMAND 1 root nginx 29 nginx nginx 30 nginx nginx 31 nginx nginx 32 nginx nginx 33 root ps
A default container on Docker 29
New for the container
mnt
own root filesystem
pid
nginx is PID 1
net
own eth0, routes, ports
ipc / uts
own IPC objects and hostname
cgroup
own cgroup appears as /
time
own clock offsets, set to 0
Shared with the host
user
UID 0 inside is UID 0 outside
kernel
one kernel, one syscall interface for every container
Namespaces change what a process can see and name. Capabilities, seccomp and AppArmor limit what it can ask the kernel to do, and cgroups limit how much it can use.

nsenter: a container's namespace with the host's tools

nsenter runs a program inside the namespaces of an existing process, and you choose which ones. That makes it the standard way to debug a container whose image has no tools: the nginx image has no ss, but entering only its network namespace runs the host's ss and ip against the container's sockets and interfaces:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ P=$(docker inspect -f '{{.State.Pid}}' lab-ns) sudo nsenter --target "$P" --net ss -tln sudo nsenter --target "$P" --net ip -br addr
State Recv-Q Send-Q Local Address:Port Peer Address:Port LISTEN 0 511 0.0.0.0:80 0.0.0.0:* LISTEN 0 511 [::]:80 [::]:* lo UNKNOWN 127.0.0.1/8 ::1/128 eth0@if1590 UP 172.17.0.3/16

nginx listens on port 80 on all addresses, and the container's eth0 holds its bridge address (yours may differ). The filesystem stayed the host's, which is why the binaries were found. Entering a different type changes what you see. With --uts the hostname is the container's; with --mount the program runs against the container's root filesystem, so head reads Alpine's /etc/os-release and the host's tools are out of reach:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ P=$(docker inspect -f '{{.State.Pid}}' lab-ns) sudo nsenter --target "$P" --uts hostname sudo nsenter --target "$P" --mount head -2 /etc/os-release
f87f04b646cd NAME="Alpine Linux" ID=alpine

nsenter needs root on the host because joining another process's namespaces requires CAP_SYS_ADMIN there. When you cannot get host root, a debug container that joins the target's namespaces is the alternative, and "Minimal bases: distroless, scratch, static and Alpine" uses that pattern for images without a shell. For the general method of tracing a connection problem, see "Network troubleshooting" in Docker in depth.

Sharing a namespace on purpose

Every type except mnt and time can be shared at docker run time. Joining another container's PID namespace is the sidecar pattern: the second container sees and can signal the first one's processes, which is how a profiler or a debugger attaches.

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm --pid container:lab-ns alpine:3.22 ps -o pid,user,comm
PID USER COMMAND 1 root nginx 29 101 nginx 30 101 nginx 31 101 nginx 32 101 nginx 41 root ps

The workers appear as user 101 here and as nginx inside the nginx container. The kernel only stores the number; the name comes from whichever /etc/passwd the viewing process reads, and the Alpine image has no nginx entry. "What root in a container really is" builds on that.

--pid host removes the PID namespace altogether. The container's pid link now carries the host's number, and the first process it sees is the host's systemd:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm --pid host alpine:3.22 sh -c 'readlink /proc/self/ns/pid; ps -o pid,comm | head -3'
pid:[4026531836] PID COMMAND 1 systemd 2 kthreadd

The setting lives in the container's run configuration, not in the image, so an image scan never sees it, and the same goes for --network host, which looks like any other network from inside. The reliable checks run from the host: compare /proc/<pid>/ns/* with /proc/1/ns/*, or read PidMode, NetworkMode, IpcMode, UTSMode and CgroupnsMode from docker inspect. What a container can do with the host's processes once it sees them depends on its capabilities and its AppArmor profile; "Host namespace sharing as attack surface" tests that.

Namespaces without Docker

unshare creates new namespaces and runs a program in them, with no daemon and no image. The classic demonstration runs it as an ordinary user and adds a user namespace with --map-root-user, so the unprivileged caller becomes root inside. On Ubuntu 24.04 and later, including this VM, that fails:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ unshare --user --map-root-user --pid --fork --mount-proc sh -c id
unshare: write failed /proc/self/uid_map: Operation not permitted
$ cat /proc/sys/kernel/apparmor_restrict_unprivileged_userns /proc/sys/kernel/unprivileged_userns_clone
1 1

The first value is kernel.apparmor_restrict_unprivileged_userns, set to 1 by Ubuntu. The second, kernel.unprivileged_userns_clone, still allows unprivileged user namespaces in principle. With the AppArmor restriction on, an unprivileged process that is not confined by a profile allowing userns loses its capabilities inside the new user namespace, so unshare cannot write the UID map and stops with the error above. Setting the sysctl to 0 would restore the old behaviour for every program on the host; the narrower fix is an AppArmor profile that grants userns to the one binary that needs it, which is what Docker's rootless setup does on Ubuntu ("Rootless Docker"). The lab VM keeps the restriction on. Other distributions without it run the command above unchanged.

As root, the same namespaces (without the user namespace) are a single command. The sh inside sets a hostname, lists processes and network links, and prints its own pid and uts links; the last line runs on the host afterwards:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ sudo unshare --pid --uts --net --ipc --mount --fork --mount-proc sh -c ' hostname box; echo "hostname inside: $(hostname)" ps -o pid,user,comm ip -br link readlink /proc/self/ns/pid /proc/self/ns/uts' echo "hostname outside: $(hostname)"
hostname inside: box PID USER COMMAND 1 root sh 4 root ps lo DOWN 00:00:00:00:00:00 <LOOPBACK> pid:[4026532448] uts:[4026532445] hostname outside: secopslog-docker

Inside, sh is PID 1 and ps sees nothing else, because --mount-proc mounted a fresh /proc for the new PID namespace. The only network interface is a loopback that is down. The hostname changed to box inside and stayed secopslog-docker outside. That is most of a container minus the image, the cgroup and the security profiles that runc adds.

Root can also create the user namespace that Docker leaves out, and map it the way user-namespace remapping does: UID 0 inside becomes UID 100000 on the host. First a world-writable scratch directory with a marker file that only host root can read, then the user namespace:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ sudo install -d -m 1777 /tmp/lab-userns sudo sh -c "echo lab-marker > /tmp/lab-userns/root-marker && chmod 600 /tmp/lab-userns/root-marker"
$ sudo unshare --user --map-users=0:100000:65536 --map-groups=0:100000:65536 --setuid 0 --setgid 0 sh -c ' id; cat /proc/self/uid_map touch /tmp/lab-userns/made-inside cat /tmp/lab-userns/root-marker' ls -ln /tmp/lab-userns
uid=0(root) gid=0(root) groups=0(root) 0 100000 65536 cat: /tmp/lab-userns/root-marker: Permission denied total 4 -rw-r--r-- 1 100000 100000 0 Oct 8 00:38 made-inside -rw------- 1 0 0 11 Oct 8 00:38 root-marker

id inside says root, and the map line reads "inside 0, outside 100000, 65536 IDs". The file created inside belongs to UID 100000 on the host, and the marker that host root owns is unreadable, because to every file outside the namespace this "root" is an unprivileged UID. Docker does not do this by default; "What root in a container really is" shows the consequence, and "User-namespace remapping" turns it on for the daemon. Clean up:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker rm -f lab-ns sudo rm -rf /tmp/lab-userns
lab-ns
Quick check
01From the host, sudo readlink /proc/1/ns/pid /proc/$P/ns/pid prints the same value twice for a container's main process $P, while the two net links differ. What do you conclude?
Incorrect — Inode numbers are unique per namespace at any moment. A matching pid link means the very same PID namespace.
Incorrect — The net links differ, so the network namespace is private. Each type is shared or private on its own.
Correct — Same pid inode as host PID 1 means it sees and numbers processes exactly as the host does.
Incorrect — That is pid_for_children. The pid link is the namespace the process itself is in.
02On the Ubuntu 26.04 lab VM, unshare --user --map-root-user --pid --fork sh as ubuntu fails with write failed /proc/self/uid_map: Operation not permitted, while sudo unshare --user ... works. Why?
Correct — kernel.apparmor_restrict_unprivileged_userns=1 strips capabilities inside a user namespace created by an unconfined unprivileged process, so the UID map cannot be written.
Incorrect — The root run created a user namespace on the same kernel, and unprivileged_userns_clone is 1.
Incorrect — unshare does not talk to Docker at all; the docker group is irrelevant to it.
Incorrect — --map-root-user maps just the caller's own UID and does not read /etc/subuid.
03A production image has no shell, no ss and no ip. You need to see which ports its process listens on without changing the image, and you have root on the host. Which command does it?
Incorrect — docker exec runs binaries from the container's filesystem; --privileged does not mount the host's /usr/bin.
Incorrect — Sockets belong to the network namespace. Entering the mount namespace switches to the image's filesystem, which has no ss.
Incorrect — Sharing the PID namespace shows processes, not sockets; that needs --network container:NAME.
Correct — Only the network namespace is entered, so the host's ss runs against the container's sockets.

Try this

Work through “Namespaces without Docker” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from namespaces from first principles, keep “Namespaces without Docker”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related