Host namespace sharing as attack surface

What --pid=host, --net=host, --ipc=host and sensitive mounts give up, and how to detect them.

Advanced11 min · lesson 19 of 24
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (2 files, 2 KB): hostns.tar.gz. The lab VM shares no folders with your computer, so fetch them inside the VM: cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-int/hostns.tar.gz && tar -xzf hostns.tar.gz, which creates ~/lab/hostns/. SHA-256: c5ce468e8aa1fadb603c81d10bee354a9ce2e0473464fe2d8bcf105ba3f517c0

Each =host flag on docker run (--pid=host, --net=host, --ipc=host, --uts=host) swaps one of a container's private namespaces for the host's, and a bind mount of a host path hands over files directly. This lesson measures what each one gives up, against a fake victim on the sec VM, and shows how to find them in docker inspect. They usually arrive through install docs: a monitoring agent asks for --pid=host --net=host, nobody looks at the flags again, and a process that later lands in that agent sees every command line on the box, the host's whole network and services bound to loopback. A non-root user and a read-only filesystem inside the container do not help, because those flags removed walls that sit underneath the container.

Watch out
Run this only in the SecOpsLog disposable lab VM (secopslog-docker-sec): it shares host namespaces and binds the host root filesystem into containers. If the VM does not exist, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec on Lima), and reset it at any point with ./setup/create-lab.sh --profile sec --recreate. On the sec VM the ubuntu account is not in the docker group, so docker runs through sudo. Every secret here is fake.

"Namespaces from first principles" covers the eight namespace types and how a normal container gets its own set. Work in ~/lab/hostns, where the lesson files go. host-setup.sh plants a safe victim view of the host: a root-owned marker file, a host process whose fake credential is on its command line, and a service bound to 127.0.0.1:9100 that serves one file from its own directory. It refuses to run outside a lab VM. The first line pulls the three images the lesson uses:

ubuntu@secopslog-docker-sec:~/lab/hostns · Docker 29.8.2
$ for i in alpine:3.22 nginx:1.30-alpine python:3.14-slim; do sudo docker pull -q $i; done sudo ./host-setup.sh up VPID=$(cat /run/lab-victim.pid); echo "victim host PID $VPID" sudo ls -l /root/lab-marker.txt curl -s http://127.0.0.1:9100/lab-index.txt
docker.io/library/alpine:3.22 docker.io/library/nginx:1.30-alpine docker.io/library/python:3.14-slim victim host PID 2100 -rw------- 1 root root 63 Oct 8 03:00 /root/lab-marker.txt host-only service, reachable on loopback

A default container shares none of this. Its namespace inode numbers differ from the host's (the kernel tags each namespace with an inode, which is the ground truth later), and its process list is its own:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ for n in pid net ipc uts cgroup; do printf "%s " $n; done; echo for n in pid net ipc uts cgroup; do readlink /proc/self/ns/$n; done echo "--- a default container:" sudo docker run --rm alpine:3.22 sh -c 'for n in pid net ipc uts cgroup; do readlink /proc/self/ns/$n; done' sudo docker run --rm alpine:3.22 ps -o pid,comm
pid net ipc uts cgroup pid:[4026531836] net:[4026531833] ipc:[4026531839] uts:[4026531838] cgroup:[4026531835] --- a default container: pid:[4026532250] net:[4026532252] ipc:[4026532249] uts:[4026532248] cgroup:[4026532251] PID COMMAND 1 ps

--pid=host puts the host process tree on screen

Share the PID namespace and the container's /proc becomes the host's. Its PID-namespace inode now matches the host's, and ps lists host daemons:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run --rm --pid=host alpine:3.22 readlink /proc/self/ns/pid sudo docker run --rm --pid=host alpine:3.22 ps -o pid,user,comm | grep -E "sleep|dockerd|sshd|PID" | head -5
pid:[4026531836] PID USER COMMAND 1076 root sshd 1274 root dockerd 1498 root sshd-session 1761 1000 sshd-session

That is already a leak, because a process's command line is readable and routinely carries secrets passed as arguments. The planted victim shows it:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ VPID=$(cat /run/lab-victim.pid) sudo docker run --rm --pid=host alpine:3.22 sh -c "tr \"\\0\" \" \" < /proc/$VPID/cmdline; echo"
db-sync --password=FAKE-HOST-ARG-SECRET 3600

The fake password on the process's argument list is right there. Reading another process's environment is a harder target, because /proc/<pid>/environ is gated by the kernel's ptrace access check. Read it with the default AppArmor profile, then with AppArmor switched off:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ VPID=$(cat /run/lab-victim.pid) sudo dmesg -C sudo docker run --rm --pid=host alpine:3.22 sh -c "cat /proc/$VPID/environ >/dev/null"; echo "default profile: exit $?" echo "AppArmor ptrace denials logged: $(sudo dmesg | grep -c "apparmor=\"DENIED\" operation=\"ptrace\"")" sudo docker run --rm --pid=host --security-opt apparmor=unconfined alpine:3.22 sh -c "cat /proc/$VPID/environ >/dev/null"; echo "apparmor=unconfined: exit $?"
cat: can't open '/proc/2100/environ': Permission denied default profile: exit 1 AppArmor ptrace denials logged: 0 cat: can't open '/proc/2100/environ': Permission denied apparmor=unconfined: exit 1

Both reads are refused, and the kernel logged no AppArmor ptrace denial, so AppArmor is not what protects the environment here. The refusal comes from the kernel's capability check: a process may read another process's memory only if it holds every capability the target holds, or holds CAP_SYS_PTRACE. The victim is a host root process with the full set; the container has the default fourteen. Adding CAP_SYS_PTRACE passes that check, and then AppArmor is the layer left. ptrace-try.py from the lesson files tries PTRACE_ATTACH on the victim process and reports what the kernel said. It refuses to run unless the lab marker file is mounted in, and it never targets PID 1:

ubuntu@secopslog-docker-sec:~/lab/hostns · Docker 29.8.2
$ VPID=$(cat /run/lab-victim.pid) sudo dmesg -C sudo docker run --rm --pid=host --cap-add SYS_PTRACE -v /etc/secopslog-lab:/etc/secopslog-lab:ro \ -v "$PWD/ptrace-try.py:/ptrace-try.py:ro" python:3.14-slim python3 /ptrace-try.py "$VPID" sudo dmesg | grep -o 'apparmor="DENIED" operation="ptrace"[^]]*profile="docker-default"[^]]*' | head -1
ptrace attach to host PID 2100 FAILED: Permission denied (errno 13) apparmor="DENIED" operation="ptrace" class="ptrace" profile="docker-default" pid=3078 comm="python3" requested_mask="trace" denied_mask="trace" peer="unconfined"

The attach fails with Permission denied, and the kernel logs apparmor="DENIED" operation="ptrace" ... profile="docker-default". Command lines still leak without any ptrace, so --pid=host on a workload that takes untrusted input is already a bad trade; the AppArmor denial is defence in depth, not a reason to allow the flag. If you also pass --security-opt apparmor=unconfined or --privileged, that last wall is gone too.

--net=host removes the network wall

Share the network namespace and the container plugs straight into the host's stack: the host's interfaces, and every service the host can reach, including the ones bound to 127.0.0.1 precisely so they stay private.

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run --rm --net=host alpine:3.22 sh -c "readlink /proc/self/ns/net; wget -qO- http://127.0.0.1:9100/lab-index.txt"
net:[4026531833] host-only service, reachable on loopback

The container's network-namespace inode matches the host's, and it reached the loopback-only service. Published-port mapping also stops meaning anything. Add -p to a --net=host container and Docker prints one warning and starts the container anyway:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run -d --name lab-nethost --net=host -p 127.0.0.1:8080:8080 nginx:1.30-alpine sudo docker port lab-nethost; echo "docker port mappings: $(sudo docker port lab-nethost | wc -l)" sudo docker rm -f lab-nethost >/dev/null
WARNING: Published ports are discarded when using host network mode e0c7419d408a98fa0e11575cc9414e122aef6978da2e21df9b66afadcbfb9abc docker port mappings: 0

The WARNING line is the only sign, and in a script or a CI log it scrolls past. docker port reports zero mappings. The service inside listens on every host interface, not the loopback address the -p suggested, and it can reach back to anything the host binds to loopback. Anywhere you find --net=host, treat every port that container opens as exposed. Rootless Docker does not change this: since Docker 29.5, rootless --net=host also joins the real host's network namespace, as "Rootless Docker" shows, although ports below 1024 still need the sysctl. Under userns-remap the daemon refuses --network=host and --privileged ("User-namespace remapping").

--ipc=host and --uts=host

--ipc=host shares System V shared memory and /dev/shm, so the container reads and writes the memory segments other processes use. A database keeping shared buffers there becomes readable:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ echo lab-fake-shm > /dev/shm/lab-marker sudo docker run --rm --ipc=host alpine:3.22 sh -c "readlink /proc/self/ns/ipc; cat /dev/shm/lab-marker" rm -f /dev/shm/lab-marker
ipc:[4026531839] lab-fake-shm

--uts=host shares the hostname and domain name. The direct damage is small, but the shared name tells an attacker which machine they are on, and a container with CAP_SYS_ADMIN could rename the host:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run --rm --uts=host alpine:3.22 hostname sudo docker run --rm alpine:3.22 hostname
secopslog-docker-sec b49dc128201e

Watch for --cgroupns=host as well, which exposes the host's control-group layout and appears in several escape write-ups. Each share is one wall the kernel stops enforcing for that container.

Sensitive mounts

A namespace share hands over a subsystem; a sensitive bind mount hands over files. Bind the host root in, even read-only, and the container reads host secrets directly:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run --rm -v /:/host:ro alpine:3.22 cat /host/root/lab-marker.txt
FAKE-HOST-SECRET (SecOpsLog lab marker, not a real credential)

It read the fake marker from the host's /root. The one mount worse than / is the Docker socket, which is full control of the daemon and so of the host; "The Docker socket and daemon hardening" covers it, and never combine it or a host-namespace share with --privileged. Other mounts to refuse by default are host /proc, /sys, /dev and devices passed with --device, each of which exposes kernel interfaces a container should not touch.

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo ls -l /var/run/docker.sock
srw-rw---- 1 root docker 0 Oct 8 03:00 /var/run/docker.sock

Detecting it

Wrapper scripts bury these flags and the person who wrote the run command has moved on, so ask the kernel. The namespace inode is proof: run readlink /proc/self/ns/pid on the host and in the container and compare.

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ HOST=$(readlink /proc/self/ns/pid) CONT=$(sudo docker run --rm --pid=host alpine:3.22 readlink /proc/self/ns/pid) ISO=$(sudo docker run --rm alpine:3.22 readlink /proc/self/ns/pid) echo "host $HOST"; echo "shared $CONT"; echo "isolated $ISO" [ "$HOST" = "$CONT" ] && echo "shared container: SAME pid namespace as host" [ "$HOST" = "$ISO" ] || echo "isolated container: different pid namespace"
host pid:[4026531836] shared pid:[4026531836] isolated pid:[4026532250] shared container: SAME pid namespace as host isolated container: different pid namespace

The shared container's inode matches the host's; the isolated one differs. For a whole fleet, docker inspect reads the same facts from the daemon config in one sweep, including the sensitive mounts:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run -d --name lab-agent --pid=host --net=host -v /:/host:ro alpine:3.22 sleep 300 >/dev/null sudo docker run -d --name lab-app alpine:3.22 sleep 300 >/dev/null sudo docker ps -q | xargs sudo docker inspect \ -f "{{.Name}} pid=[{{.HostConfig.PidMode}}] net=[{{.HostConfig.NetworkMode}}] ipc=[{{.HostConfig.IpcMode}}] uts=[{{.HostConfig.UTSMode}}] mounts=[{{range .Mounts}}{{.Source}}:{{.Destination}} {{end}}]"
/lab-app pid=[] net=[bridge] ipc=[private] uts=[] mounts=[] /lab-agent pid=[host] net=[host] ipc=[private] uts=[] mounts=[/:/host ]

lab-agent shows pid=host net=host and a /:/host mount, exactly the shape to review now; lab-app shares nothing. Put this sweep on a schedule and keep a short written list of the exceptions you have granted, because an undocumented one quietly becomes permanent. A pipeline step that greps Compose files and run scripts for pid=host, network_mode: host, hostPID and hostNetwork catches most of these before they reach a host.

When a tool genuinely needs one

Treat these flags like --privileged: off by default, granted only when a tool has proved it needs one specific namespace and cannot do the job otherwise, never on a workload that takes untrusted input. When you grant one, grant that one alone and keep the container boring: drop capabilities, read-only rootfs, no new privileges, and publish on loopback rather than sharing the network.

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run -d --name lab-nodeexp --pid=host -p 127.0.0.1:9101:9101 \ --cap-drop ALL --read-only --security-opt no-new-privileges=true alpine:3.22 sleep 300 >/dev/null sudo docker inspect lab-nodeexp -f "pid={{.HostConfig.PidMode}} net={{.HostConfig.NetworkMode}} caps_dropped={{.HostConfig.CapDrop}} nnp={{index .HostConfig.SecurityOpt 0}}"
pid=host net=bridge caps_dropped=[ALL] nnp=no-new-privileges=true

This agent gets --pid=host and nothing else: net=bridge, all capabilities dropped, no-new-privileges set, its port on 127.0.0.1. Kubernetes calls the same three controls hostPID, hostNetwork and hostIPC, and both its Baseline and Restricted Pod Security levels reject them, so admission control can make the argument for you.

ubuntu@secopslog-docker-sec:~/lab/hostns · Docker 29.8.2
$ sudo docker rm -f lab-agent lab-app lab-nodeexp >/dev/null sudo ./host-setup.sh down
Quick check
01A container runs as a non-root user, drops every capability, and mounts its root filesystem read-only. Which single run flag can still expose the host's secrets to an attacker inside it?
Incorrect — Read-only only blocks writes to the container's own filesystem; it shares nothing with the host.
Correct — Command lines are readable without any capability and routinely carry tokens and connection strings; the container's own hardening sits above a wall this flag removed.
Incorrect — Dropping capabilities is part of the hardening and does not remove seccomp or open a window onto the host.
Incorrect — Running as an unprivileged UID narrows what the process can do and opens no path to the host.
02You want to prove whether a running container shares the host's PID namespace, without trusting the wrapper script that launched it. What settles it?
Incorrect — Process counts vary for many reasons and are never proof of a shared namespace.
Incorrect — The base image is unrelated; sharing is decided at run time by the flags passed.
Incorrect — Sharing comes from the run flags regardless of the in-container user; a non-root container shares the host PID namespace happily.
Correct — The kernel tags each namespace with an inode, so a matching number means the same namespace.
03Someone starts a container with --net=host and adds -p 127.0.0.1:8080:8080 to "keep it private". What actually happens to the service on 8080?
Incorrect — With --net=host the -p mapping is discarded, so the loopback binding you intended never happens.
Incorrect — Docker accepts it with only a warning line, which is why it slips past review.
Correct — Host networking makes published-port rules meaningless, so the port is fully exposed, and the container also reaches host loopback services.
Incorrect — The service still listens; it listens on every interface instead of loopback, which is more exposed, not less.

Try this

Work through “When a tool genuinely needs one” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from host namespace sharing as attack surface, keep “When a tool genuinely needs one”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related