Namespaces from first principles
What each namespace isolates and how to see it with lsns and nsenter.
Run sudo lsns against the main process of an ordinary nginx container and you get eight lines. Seven name a namespace that exists only for that container. The eighth, the user namespace, carries the same number as the one systemd runs in. Most of what this course does starts from that one shared line, and from the fact that the kernel has no container object at all: there are processes, and there are namespaces, cgroups, capabilities and security profiles attached to them.
This course, Advanced container security, assumes Docker for beginners and the image, BuildKit, engine, storage, networking and production/CI parts of Docker in depth, in particular "Engine architecture on Docker 29" and "Networks, drivers and DNS". It does not assume Swarm or the recipes. Commands run on the main lab VM, secopslog-docker, built from the lab kit as "Installing Docker Engine and running your first container" in Docker for beginners describes (./setup/create-lab.sh, then multipass shell secopslog-docker). Lessons that change the daemon or the host, such as rootless Docker, user-namespace remapping, AppArmor profiles and the escape demonstrations, use the second VM, secopslog-docker-sec (./setup/create-lab.sh --profile sec; rebuild it with --recreate when you want it clean), and say so in their first lines. Host-side inspection uses sudo inside the VM, where the ubuntu account has it without a password. Each lesson that has files shows a Lesson files box with the command that fetches and unpacks them inside the VM, into ~/lab/<lesson>. This lesson needs none and changes nothing on the host.
Eight namespace types
A namespace wraps one kind of kernel resource so that the processes inside it get their own instance: their own process ID numbers, their own mount table, their own network stack. Every process belongs to exactly one namespace of each type, and /proc/<pid>/ns/ has a link per type. Your login shell on the VM is in the host's initial set:
The number in brackets is the inode number of the namespace on the kernel's internal nsfs filesystem. Same number, same namespace; a different number is a different namespace. That comparison is the whole detection method in this lesson. The initial namespaces have fixed numbers in the kernel source, so pid:[4026531836] reads the same on every Linux host. On recent kernels, including this VM's 7.0, that holds for all eight, mount (4026531832) and network (4026531833) included; older kernels allocate those two at boot, so they vary there. pid_for_children and time_for_children exist because those two types take effect for children: a process that calls unshare for them moves its future children, not itself.
What a default container gets
Start one nginx container and look at it from the host. docker inspect gives you its main process ID on the host, which is what every host-side tool wants:
lsns -p lists the namespaces of one process. Seven rows belong to the nginx master process (five processes share them: the master and four workers). The user row lists systemd as the first process in it, because the container never got a user namespace of its own. Your process IDs and namespace numbers will differ; the pattern will not. The same comparison as a loop over all eight types, host PID 1 against the container:
Seven numbers differ and user matches. Docker 29.8.2 with runc 1.5.1 also gives each container its own time namespace; it does not change the clocks (both offsets are zero), so you will rarely notice it:
Docker records the choices in the container's HostConfig. Empty PidMode and UTSMode mean private; ipc=private is the daemon's default IPC mode; cgroupns=private is the default on cgroup v2 hosts; an empty UsernsMode means the host user namespace. These fields are what an audit script reads, since they cannot be faked from inside the container:
From the inside, the private namespaces show up as a short hostname (the container ID), a cgroup path of 0::/ (the cgroup namespace makes the container's own cgroup look like the root of the tree) and a process list that starts with nginx as PID 1:
nsenter: a container's namespace with the host's tools
nsenter runs a program inside the namespaces of an existing process, and you choose which ones. That makes it the standard way to debug a container whose image has no tools: the nginx image has no ss, but entering only its network namespace runs the host's ss and ip against the container's sockets and interfaces:
nginx listens on port 80 on all addresses, and the container's eth0 holds its bridge address (yours may differ). The filesystem stayed the host's, which is why the binaries were found. Entering a different type changes what you see. With --uts the hostname is the container's; with --mount the program runs against the container's root filesystem, so head reads Alpine's /etc/os-release and the host's tools are out of reach:
nsenter needs root on the host because joining another process's namespaces requires CAP_SYS_ADMIN there. When you cannot get host root, a debug container that joins the target's namespaces is the alternative, and "Minimal bases: distroless, scratch, static and Alpine" uses that pattern for images without a shell. For the general method of tracing a connection problem, see "Network troubleshooting" in Docker in depth.
Sharing a namespace on purpose
Every type except mnt and time can be shared at docker run time. Joining another container's PID namespace is the sidecar pattern: the second container sees and can signal the first one's processes, which is how a profiler or a debugger attaches.
The workers appear as user 101 here and as nginx inside the nginx container. The kernel only stores the number; the name comes from whichever /etc/passwd the viewing process reads, and the Alpine image has no nginx entry. "What root in a container really is" builds on that.
--pid host removes the PID namespace altogether. The container's pid link now carries the host's number, and the first process it sees is the host's systemd:
The setting lives in the container's run configuration, not in the image, so an image scan never sees it, and the same goes for --network host, which looks like any other network from inside. The reliable checks run from the host: compare /proc/<pid>/ns/* with /proc/1/ns/*, or read PidMode, NetworkMode, IpcMode, UTSMode and CgroupnsMode from docker inspect. What a container can do with the host's processes once it sees them depends on its capabilities and its AppArmor profile; "Host namespace sharing as attack surface" tests that.
Namespaces without Docker
unshare creates new namespaces and runs a program in them, with no daemon and no image. The classic demonstration runs it as an ordinary user and adds a user namespace with --map-root-user, so the unprivileged caller becomes root inside. On Ubuntu 24.04 and later, including this VM, that fails:
The first value is kernel.apparmor_restrict_unprivileged_userns, set to 1 by Ubuntu. The second, kernel.unprivileged_userns_clone, still allows unprivileged user namespaces in principle. With the AppArmor restriction on, an unprivileged process that is not confined by a profile allowing userns loses its capabilities inside the new user namespace, so unshare cannot write the UID map and stops with the error above. Setting the sysctl to 0 would restore the old behaviour for every program on the host; the narrower fix is an AppArmor profile that grants userns to the one binary that needs it, which is what Docker's rootless setup does on Ubuntu ("Rootless Docker"). The lab VM keeps the restriction on. Other distributions without it run the command above unchanged.
As root, the same namespaces (without the user namespace) are a single command. The sh inside sets a hostname, lists processes and network links, and prints its own pid and uts links; the last line runs on the host afterwards:
Inside, sh is PID 1 and ps sees nothing else, because --mount-proc mounted a fresh /proc for the new PID namespace. The only network interface is a loopback that is down. The hostname changed to box inside and stayed secopslog-docker outside. That is most of a container minus the image, the cgroup and the security profiles that runc adds.
Root can also create the user namespace that Docker leaves out, and map it the way user-namespace remapping does: UID 0 inside becomes UID 100000 on the host. First a world-writable scratch directory with a marker file that only host root can read, then the user namespace:
id inside says root, and the map line reads "inside 0, outside 100000, 65536 IDs". The file created inside belongs to UID 100000 on the host, and the marker that host root owns is unreadable, because to every file outside the namespace this "root" is an unprivileged UID. Docker does not do this by default; "What root in a container really is" shows the consequence, and "User-namespace remapping" turns it on for the daemon. Clean up:
sudo readlink /proc/1/ns/pid /proc/$P/ns/pid prints the same value twice for a container's main process $P, while the two net links differ. What do you conclude?unshare --user --map-root-user --pid --fork sh as ubuntu fails with write failed /proc/self/uid_map: Operation not permitted, while sudo unshare --user ... works. Why?ss and no ip. You need to see which ports its process listens on without changing the image, and you have root on the host. Which command does it?Try this
Work through “Namespaces without Docker” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from namespaces from first principles, keep “Namespaces without Docker”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.