Host namespace sharing as attack surface
What --pid=host, --net=host, --ipc=host and sensitive mounts give up, and how to detect them.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-int/hostns.tar.gz && tar -xzf hostns.tar.gz, which creates ~/lab/hostns/. SHA-256: c5ce468e8aa1fadb603c81d10bee354a9ce2e0473464fe2d8bcf105ba3f517c0Each =host flag on docker run (--pid=host, --net=host, --ipc=host, --uts=host) swaps one of a container's private namespaces for the host's, and a bind mount of a host path hands over files directly. This lesson measures what each one gives up, against a fake victim on the sec VM, and shows how to find them in docker inspect. They usually arrive through install docs: a monitoring agent asks for --pid=host --net=host, nobody looks at the flags again, and a process that later lands in that agent sees every command line on the box, the host's whole network and services bound to loopback. A non-root user and a read-only filesystem inside the container do not help, because those flags removed walls that sit underneath the container.
secopslog-docker-sec): it shares host namespaces and binds the host root filesystem into containers. If the VM does not exist, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec on Lima), and reset it at any point with ./setup/create-lab.sh --profile sec --recreate. On the sec VM the ubuntu account is not in the docker group, so docker runs through sudo. Every secret here is fake."Namespaces from first principles" covers the eight namespace types and how a normal container gets its own set. Work in ~/lab/hostns, where the lesson files go. host-setup.sh plants a safe victim view of the host: a root-owned marker file, a host process whose fake credential is on its command line, and a service bound to 127.0.0.1:9100 that serves one file from its own directory. It refuses to run outside a lab VM. The first line pulls the three images the lesson uses:
A default container shares none of this. Its namespace inode numbers differ from the host's (the kernel tags each namespace with an inode, which is the ground truth later), and its process list is its own:
--pid=host puts the host process tree on screen
Share the PID namespace and the container's /proc becomes the host's. Its PID-namespace inode now matches the host's, and ps lists host daemons:
That is already a leak, because a process's command line is readable and routinely carries secrets passed as arguments. The planted victim shows it:
The fake password on the process's argument list is right there. Reading another process's environment is a harder target, because /proc/<pid>/environ is gated by the kernel's ptrace access check. Read it with the default AppArmor profile, then with AppArmor switched off:
Both reads are refused, and the kernel logged no AppArmor ptrace denial, so AppArmor is not what protects the environment here. The refusal comes from the kernel's capability check: a process may read another process's memory only if it holds every capability the target holds, or holds CAP_SYS_PTRACE. The victim is a host root process with the full set; the container has the default fourteen. Adding CAP_SYS_PTRACE passes that check, and then AppArmor is the layer left. ptrace-try.py from the lesson files tries PTRACE_ATTACH on the victim process and reports what the kernel said. It refuses to run unless the lab marker file is mounted in, and it never targets PID 1:
The attach fails with Permission denied, and the kernel logs apparmor="DENIED" operation="ptrace" ... profile="docker-default". Command lines still leak without any ptrace, so --pid=host on a workload that takes untrusted input is already a bad trade; the AppArmor denial is defence in depth, not a reason to allow the flag. If you also pass --security-opt apparmor=unconfined or --privileged, that last wall is gone too.
--net=host removes the network wall
Share the network namespace and the container plugs straight into the host's stack: the host's interfaces, and every service the host can reach, including the ones bound to 127.0.0.1 precisely so they stay private.
The container's network-namespace inode matches the host's, and it reached the loopback-only service. Published-port mapping also stops meaning anything. Add -p to a --net=host container and Docker prints one warning and starts the container anyway:
The WARNING line is the only sign, and in a script or a CI log it scrolls past. docker port reports zero mappings. The service inside listens on every host interface, not the loopback address the -p suggested, and it can reach back to anything the host binds to loopback. Anywhere you find --net=host, treat every port that container opens as exposed. Rootless Docker does not change this: since Docker 29.5, rootless --net=host also joins the real host's network namespace, as "Rootless Docker" shows, although ports below 1024 still need the sysctl. Under userns-remap the daemon refuses --network=host and --privileged ("User-namespace remapping").
--ipc=host and --uts=host
--ipc=host shares System V shared memory and /dev/shm, so the container reads and writes the memory segments other processes use. A database keeping shared buffers there becomes readable:
--uts=host shares the hostname and domain name. The direct damage is small, but the shared name tells an attacker which machine they are on, and a container with CAP_SYS_ADMIN could rename the host:
Watch for --cgroupns=host as well, which exposes the host's control-group layout and appears in several escape write-ups. Each share is one wall the kernel stops enforcing for that container.
Sensitive mounts
A namespace share hands over a subsystem; a sensitive bind mount hands over files. Bind the host root in, even read-only, and the container reads host secrets directly:
It read the fake marker from the host's /root. The one mount worse than / is the Docker socket, which is full control of the daemon and so of the host; "The Docker socket and daemon hardening" covers it, and never combine it or a host-namespace share with --privileged. Other mounts to refuse by default are host /proc, /sys, /dev and devices passed with --device, each of which exposes kernel interfaces a container should not touch.
Detecting it
Wrapper scripts bury these flags and the person who wrote the run command has moved on, so ask the kernel. The namespace inode is proof: run readlink /proc/self/ns/pid on the host and in the container and compare.
The shared container's inode matches the host's; the isolated one differs. For a whole fleet, docker inspect reads the same facts from the daemon config in one sweep, including the sensitive mounts:
lab-agent shows pid=host net=host and a /:/host mount, exactly the shape to review now; lab-app shares nothing. Put this sweep on a schedule and keep a short written list of the exceptions you have granted, because an undocumented one quietly becomes permanent. A pipeline step that greps Compose files and run scripts for pid=host, network_mode: host, hostPID and hostNetwork catches most of these before they reach a host.
When a tool genuinely needs one
Treat these flags like --privileged: off by default, granted only when a tool has proved it needs one specific namespace and cannot do the job otherwise, never on a workload that takes untrusted input. When you grant one, grant that one alone and keep the container boring: drop capabilities, read-only rootfs, no new privileges, and publish on loopback rather than sharing the network.
This agent gets --pid=host and nothing else: net=bridge, all capabilities dropped, no-new-privileges set, its port on 127.0.0.1. Kubernetes calls the same three controls hostPID, hostNetwork and hostIPC, and both its Baseline and Restricted Pod Security levels reject them, so admission control can make the argument for you.
--net=host and adds -p 127.0.0.1:8080:8080 to "keep it private". What actually happens to the service on 8080?--net=host the -p mapping is discarded, so the loopback binding you intended never happens.Try this
Work through “When a tool genuinely needs one” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from host namespace sharing as attack surface, keep “When a tool genuinely needs one”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.