Containers and namespaces from the host
Map container processes and spot escapes.
A container is not a small virtual machine. It is one or more ordinary Linux processes that the kernel has placed in their own namespaces (private views of the process table, mounts, network and so on) and a cgroup (a resource-accounting group). That is the defender's advantage: from the host, a container is just processes, visible in ps and /proc, and everything that isolates them is readable through interfaces you already use. This lesson shows how to find a container's processes from the host whatever started them, map a container PID to a host PID, read and enter the namespaces that fence it in, and tell an ordinary container from one that can reach the host, the other road from a foothold to root. The lab uses podman (install it with sudo apt install podman on Ubuntu; it is RHEL's native engine) with a small pinned image; nothing is escaped, only inspected. linux-perf/k-ns teaches the mechanism; here the angle is detection.
A container is only processes
Pull a small image and start a container that runs sleep 900 with no network (--network none). Both commands print an ID: the image's, then the new container's.
Now find it from the host. The engine records the host-side PID of the container's first process, and ps shows it like any other.
The sleep runs as host PID 309726, owned by root, and its parent is conmon, a small monitor that podman attaches to each container to hold its output and report its exit. conmon's parent is PID 1, so the container outlives the command that launched it. The runtime that built the namespaces, crun, has already exited. On the host a container is a normal process tree; the tell is not the process but where it sits, which the rest of this lesson reads.
Inside the container that same process thinks it is PID 1. Two views make the mapping explicit.
podman top ... hpid pid prints the host PID (309726) beside the in-container PID (1). NSpid: 309726 1 in /proc/PID/status says the same: the process's PID in each nested PID namespace, outermost first. An alert that names an in-container PID means nothing on the host until you translate it this way; and for any suspicious host PID, a second number on its NSpid line says it runs inside a PID namespace.
What isolates it: namespaces and the cgroup
The container's process is in its own namespaces. Compare each one against the host's PID 1 to see which views it has been given privately.
Six namespaces differ from the host (mnt, pid, net, ipc, uts, cgroup): the container has its own mounts, process table, network stack, IPC objects, hostname and cgroup view. user does not differ. The container shares the host's user namespace, so its root is the host's real UID 0, held back only by the dropped capabilities and the access label shown below, not by a UID mapping: one missing capability or one bad mount away from root on the host. A rootless container (podman run by an ordinary user) would show a private user namespace mapping container-root to an unprivileged host UID.
The cgroup line says who manages it, and lsns counts the processes sharing each namespace.
The cgroup path /machine.slice/libpod-<id>.scope/container marks this as a podman-managed container (systemd puts machines and containers under machine.slice), and the long id ties the process back to podman inspect. lsns shows each of the six namespaces holding a single process, while the time and user namespaces are the host's own (NPROCS counts the host processes in each at that moment, about 120 here; the counts move as processes start and exit). The cgroup path is the stronger container marker of the two, as the hunt at the end of this lesson shows.
Because these are just kernel objects, you can enter them from the host with nsenter, which is how you inspect a container without the engine or a shell inside it.
Entering the network namespace (-n) shows only a loopback interface, because this container was run with --network none. Entering the mount namespace (-m) drops you into the container's root filesystem: NAME="Alpine Linux" and a root that is the image's, not the host's Ubuntu. nsenter against the host PID is a defender's tool: it lets you list files, sockets or processes as the container sees them during an investigation, using only host privileges.
What it may do: capabilities and the access label
A rootful container's root is restrained by two things. The first is a reduced capability set: the kernel splits root's power into separate capabilities (linux-det/peloc), and podman keeps only a small default subset. The second is a mandatory-access-control label: an AppArmor profile on Ubuntu, an SELinux type on RHEL.
CapEff: 00000000800405fb decodes to eleven capabilities, not the full forty-plus. cap_sys_admin, cap_sys_module, cap_sys_ptrace and the other broad ones are absent, which is why a default container cannot load a kernel module, re-mount the host filesystem or trace host processes even though it runs as UID 0. The AppArmor label containers-default-0.66.0 (enforce) is podman's default profile, confining the process further. On RHEL the confinement is SELinux, and the label shows in ps -eZ.
The type container_t with a unique category pair (s0:c788,c802) is how SELinux isolates one container from another and from the host: two containers get different categories, so even as root neither can touch the other's files. The capability set is the same reduced eleven. So on both platforms the defensive baseline for a container process is: a known-small capability set and a confining label. A container process with the full capability set, or an unconfined label, has neither guardrail, and that is the next thing to recognise.
A privileged container, and the escape signals
A privileged container is one started with --privileged, often with extra host access bolted on. It is sometimes legitimate (a monitoring agent, a storage driver), but it is also what an attacker asks for, because it removes the guardrails above. Start one the way it is usually found, with the host's process namespace (--pid=host) and the host's root filesystem mounted at /host (-v /:/host), and read it from the host without entering it.
podman inspect reports privileged=true, pidmode=host, and the host root mounted at /host. The process carries CapEff: 000001ffffffffff, the full capability set, and its label is crun (unconfined): Ubuntu ships a crun AppArmor profile, but it is declared unconfined, so no AppArmor restriction applies.
Those three facts, full capabilities, no confining label, host resources mapped in, are the signature of a container that can trivially become the host. The namespace view confirms it.
--pid=host means the container shares the host's PID namespace (pid matches PID 1's), so its processes can see and signal every process on the host. findmnt against its PID shows the host's real root filesystem (/dev/vda1) mounted inside it at /host, read-write (rw): from there, root in the container reads and writes every file on the host. (A :ro suffix on -v would mount it read-only, but a privileged container's root holds CAP_SYS_ADMIN and can remount it, so treat any host root mount in a privileged container as writable.) You do not demonstrate the escape; you recognise the shape.
machine.slice/libpod-*, docker-*, cri-containerd-*, kubepods) or its PID namespace differs from PID 1's. On the host every root process has a full CapEff, and on Ubuntu most are AppArmor unconfined, so an unscoped rule fires on the whole process table. Inside that scope, three findings mark a container that can reach the host: a full capability set or an unconfined label; the host's PID or network namespace (--pid=host, --net=host: its /proc/PID/ns/* matches PID 1's); and a host path mounted in, above all /, /proc, /dev or an engine socket (/var/run/docker.sock, /run/podman/podman.sock, /run/containerd/containerd.sock). Baseline the containers that legitimately need these, and alert on any other container that has them.Finding every container, whatever started it
The engine's own list is the obvious starting point, and it is incomplete.
sudo podman ps lists the containers of root's podman and nothing else. Rootless podman keeps each user's containers in that user's own storage, so root's list never shows them (sudo -iu name podman ps does, per account). Docker, containerd and CRI-O (the Kubernetes runtimes, queried with docker ps and crictl ps) keep their own lists. And a process can enter private namespaces with unshare without any engine at all. Start one of those as a stand-in for a sandbox no engine knows about.
Now ask the kernel instead. lsns -t pid lists every PID namespace on the host with its lowest-numbered process, whoever created it.
Three PID namespaces: the host's, ct-web's, and the unshare sandbox podman did not list. Their cgroups tell them apart: ct-web sits in a machine.slice/libpod-* scope, the sandbox in a login session's scope (here user 502's, the account the lab VM's tool runs commands through; on your host, your own session). A private namespace whose cgroup is neither a container scope nor a service you can explain is the finding to chase. ct-priv is missing because --pid=host keeps it in the host's PID namespace, so also read every process's cgroup for container scopes.
Both podman containers show up, one scope each, including the privileged one the PID-namespace view missed. The same pattern names Docker (docker-), containerd (cri-containerd-) and CRI-O (crio-) containers. A private mount namespace on its own proves nothing, though.
systemd gives services their own mount namespace when a unit uses sandboxing settings such as PrivateTmp= or ProtectSystem=, so resolved, udevd, networkd, ModemManager, polkitd, logind, fwupd and chrony all have one on this server, none of them in a container. PrivateUsers= and PrivateNetwork= do the same for user and network namespaces. So never alert on a private namespace alone: pair it with the cgroup, where a system.slice unit whose sandboxing explains it is expected and a login session or an unknown scope is not, and audit unshare and setns calls from accounts with no reason to make them. The Rocky lab shows both sides.
irqbalance has a user namespace of its own because its unit sets PrivateUsers=true: expected. Root's podman lists only ct-web, yet lsns -t user also finds a user namespace owned by lima running a rootless containerd. That one belongs to the lab VM's tool, not to RHEL, but an intruder's rootless container looks exactly the same: a non-root owner, a private user namespace, and no entry in root's engine.
Ubuntu restricts unprivileged user namespaces
User namespaces let an unprivileged user become root inside a namespace of their own, which is what makes rootless containers possible. The same feature has been the way into many kernel exploits, because it hands an ordinary user code paths that expect privilege. Ubuntu 26.04 ships a mitigation on by default; linux-hard/sysctlkern showed the switch, and here is exactly what it does.
kernel.apparmor_restrict_unprivileged_userns is 1. An unprivileged user can still create a user namespace (unshare -U succeeds, leaving them mapped to nobody; every ID with no mapping shows as the overflow ID 65534, so each of deploy's supplementary groups appears as nogroup), but the privileged step, mapping their UID to root inside it with -r, is refused. The restriction is enforced through AppArmor, so the denial is recorded.
The audit record names the mechanism: AppArmor DENIED the CAP_SYS_ADMIN the mapping needs, under a profile called unprivileged_userns that the kernel attaches to any unconfined program trying this. (auditd is installed on this lab host; linux-det/auditpipe covers reading these records, and without auditd the same line lands in journalctl -k. capability=21 is CAP_SYS_ADMIN on every architecture: capability numbers, unlike system-call numbers, do not vary.) Programs that genuinely need user namespaces are allowed through a profile that grants the userns permission, which is how podman still works.
The podman profile is flags=(unconfined) and lists userns,, so podman may create user namespaces while a bare unshare from an ordinary account may not. RHEL takes a different path. The sysctl does not exist there, user.max_user_namespaces is sized from memory rather than set to zero, and unconfined_t does not restrict user namespaces, so the same unshare -U -r succeeds on the Rocky lab.
On RHEL, then, every local account can reach the kernel code that user namespaces expose. Where no rootless containers or other user-namespace users run, setting user.max_user_namespaces = 0 in a sysctl.d file closes it (linux-hard/sysctlkern); where they do, audit unshare and setns by accounts that have no reason to call them. Know which platform you are on before you conclude that an unprivileged user namespace is or is not a concern.
Try this
On the Ubuntu lab, with podman installed, start sudo podman run -d --name t --network none docker.io/library/alpine:3.23.4 sleep 600. Find its host PID with sudo podman inspect -f '{{.State.Pid}}' t and confirm the mapping with sudo grep NSpid /proc/<pid>/status (host PID, then 1). Compare namespaces with for ns in pid net user; do echo $ns $(sudo readlink /proc/1/ns/$ns) $(sudo readlink /proc/<pid>/ns/$ns); done: pid and net differ, user matches. Run the cgroup-scope count from this lesson and find one more libpod- scope. Decode CapEff with sudo capsh --decode=<value> and see the eleven capabilities, then remove it with sudo podman rm -f t and confirm the scope count drops again.
Takeaway
Find containers from the kernel side (container cgroup scopes and private PID or user namespaces) before you ask any engine, then treat each as processes: map its host PID and read its namespaces, capabilities and LSM label from /proc. Alert on the escape signals only inside container scopes, and baseline the containers that are allowed to have them.
ps on the host and see no miner at PID 1 (it is systemd). What has gone wrong, and what do you do?podman top ... hpid pid or /proc/PID/status NSpid, then act on the host PID./proc/PID/ns/*, you find every namespace differs from the host's PID 1 except user, which is identical. Why does that matter for containment?/, is a direct route out; combined with a rootful container's UID 0 it is full read/write of the host.