Container forensics and incident response

Contain, capture and diff a suspect container before you destroy it.

Advanced14 min · lesson 20 of 24

The reflex that ends an outage, kill it and let it come back clean, destroys the evidence of a compromise. A container keeps that evidence in places the quick fixes erase in seconds: live process memory and open sockets (gone when the process dies), the container's log and configuration (gone on docker rm with the default json-file driver), and the writable layer stacked on the read-only image (also gone on docker rm). The job in the first few minutes is preservation, in one fixed order: contain, capture, investigate, remediate. This lesson runs that order on a benign suspect on the main lab VM (secopslog-docker), in ~/lab/forensics. The suspect plants its own fake evidence and exploits nothing.

The order that keeps the evidence readable
1Contain
pause, record PID, config, logs and sockets, then cut the network
2Capture
host /proc, docker cp, diff, export, commit, save, hashes
3Investigate
line up the artifacts against the host timeline
4Remediate
neutralise restart policy, rm, rebuild from a clean image
Jumping to remediate redeploys the same vulnerable image an hour later, having learned nothing about the way in.

The suspect logs one line, drops a copy of a binary, adds a crontab line, opens a TCP connection to lab-c2 (a stand-in for an attacker's server on the same network) and then sleeps, with the fake command-and-control address in its environment:

ubuntu@secopslog-docker:~/lab/forensics · Docker 29.8.2
$ docker pull -q alpine:3.22 docker network create lab-evnet >/dev/null docker run -d --name lab-c2 --network lab-evnet alpine:3.22 sh -c "sleep 3600 | nc -l -p 4444 >/dev/null" >/dev/null sleep 1 docker run -d --name lab-suspect --network lab-evnet --restart unless-stopped \ -e C2_HOST=lab-c2:4444 alpine:3.22 sh -c ' echo "worker started" cp /bin/busybox /tmp/.miner printf "* * * * * root /tmp/.miner\n" >> /etc/crontab (sleep 3600 | nc lab-c2 4444 >/dev/null &) exec sleep 3600' >/dev/null sleep 2; docker ps --filter name=lab- --format "{{.Names}} {{.Status}}"
Error response from daemon: failed to resolve reference "docker.io/library/alpine:3.22": failed to do request: Head "https://registry-1.docker.io/v2/library/alpine/manifests/3.22": dial tcp: lookup registry-1.docker.io on 127.0.0.53:53: read udp 127.0.0.1:58537->127.0.0.53:53: i/o timeout lab-suspect Up 2 seconds lab-c2 Up 3 seconds

Contain: freeze, record, then cut the network

Freeze first. docker pause stops every thread at once through the cgroup freezer. It acts underneath the process, so there is no signal for an attacker to trap and nothing gets a chance to run a clean-up, and a paused container refuses docker exec, which is why every capture step runs from the host:

ubuntu@secopslog-docker:~/lab/forensics · Docker 29.8.2
$ docker pause lab-suspect docker inspect -f "{{.State.Status}}" lab-suspect docker exec lab-suspect true 2>&1; echo "exec while paused: exit $?"
lab-suspect paused Error response from daemon: Container lab-suspect is paused, unpause the container before exec exec while paused: exit 1

Then write down what the daemon knows before anything changes it. The host-side PID is the number the kernel gives the container's main process (not the 1 it sees inside), and the next steps read the kernel's view of exactly that process. docker inspect is the full record of how the container was started: image digest, mounts, environment, restart policy, labels and network settings. docker logs is everything the process wrote to stdout and stderr. Both disappear with docker rm, so they go into the evidence directory now. The directory is mode 700; on a real host, write to a root-only location on separate storage and copy it off the host as soon as you can:

ubuntu@secopslog-docker:~/lab/forensics · Docker 29.8.2
$ install -d -m 700 ir-42 PID=$(docker inspect -f "{{.State.Pid}}" lab-suspect); echo "host PID $PID" | tee ir-42/pid.txt docker inspect lab-suspect > ir-42/inspect.json docker logs --timestamps lab-suspect > ir-42/logs.txt 2>&1 cat ir-42/logs.txt ls -l ir-42
host PID 684837 2026-10-07T22:57:04.163959142Z worker started total 16 -rw-r--r-- 1 ubuntu ubuntu 7892 Oct 8 04:27 inspect.json -rw-r--r-- 1 ubuntu ubuntu 46 Oct 8 04:27 logs.txt -rw-r--r-- 1 ubuntu ubuntu 16 Oct 8 04:27 pid.txt

Record the connections next, from the host, with nsenter joining only the container's network namespace ("Namespaces from first principles" introduced the method), so the ss binary is the host's and cannot have been replaced by the attacker:

ubuntu@secopslog-docker:~/lab/forensics · Docker 29.8.2
$ PID=$(cut -d" " -f3 ir-42/pid.txt) sudo nsenter -t "$PID" -n ss -tanp | tee ir-42/sockets.txt
State Recv-Q Send-Q Local Address:Port Peer Address:PortProcess LISTEN 0 4096 127.0.0.11:43825 0.0.0.0:* users:(("dockerd",pid=17027,fd=44)) ESTAB 0 0 172.18.0.3:37913 172.18.0.2:4444 users:(("nc",pid=684873,fd=3))

The ESTAB line is the open channel: the suspect's nc connected to 172.18.0.2:4444, the address of lab-c2. On a real incident that peer address is often the most valuable volatile fact you collect. (The LISTEN line on 127.0.0.11 is Docker's embedded DNS resolver.) Now cut the network:

ubuntu@secopslog-docker:~/lab/forensics · Docker 29.8.2
$ docker network disconnect lab-evnet lab-suspect docker inspect -f "attached networks: {{len .NetworkSettings.Networks}}" lab-suspect
attached networks: 0

Disconnecting removes the container's interface, so the connection can no longer carry traffic: anything sent on it fails and the socket times out, and the address that tied it to the network is gone. That is why the socket table was saved first. If the container is actively sending data out, cut first and accept the loss, or drop its traffic with a host firewall rule on its address, which stops the flow and leaves the socket state readable.

Capture from the host, not from inside

A compromised container controls its own userspace, so its ps, ls and cat can lie, and its own files can be edited. Read the kernel's view from the host instead, and know which parts of it you can trust. Kernel-maintained entries such as /proc/<pid>/exe, fd, maps and status cannot be forged from inside: exe is the running binary, and it still resolves even if the attacker unlinked the file (it then reads (deleted)). /proc/<pid>/cmdline and environ are different. The kernel reads them out of the process's own memory, so a process can overwrite them (argument rewriting is routine malware behaviour), and environ shows the environment block the process started with, not variables it set later. Treat those two as claims to corroborate. Pull the dropped artifact with docker cp, which reads the filesystem through the daemon and works on a paused container:

ubuntu@secopslog-docker:~/lab/forensics · Docker 29.8.2
$ PID=$(cut -d" " -f3 ir-42/pid.txt) sudo ls -l /proc/$PID/exe
lrwxrwxrwx 1 root root 0 Oct 8 04:27 /proc/684837/exe -> /bin/busybox
$ docker cp lab-suspect:/tmp/.miner ir-42/miner.bin file ir-42/miner.bin | cut -d, -f1
Successfully copied 919kB (transferred 921kB) to /home/ubuntu/lab/forensics/ir-42/miner.bin ir-42/miner.bin: ELF 64-bit LSB pie executable
$ PID=$(cut -d" " -f3 ir-42/pid.txt) sudo cat /proc/$PID/environ | tr "\0" "\n" > ir-42/environ.txt grep C2 ir-42/environ.txt
C2_HOST=lab-c2:4444

The live binary resolves, the dropped file comes out as an ELF executable, and the fake C2 address is read from the process's environment block.

Diff, export, commit, hash

The image ships read-only and never changes, so everything the container wrote sits in the thin writable layer on top. docker diff reads that layer directly, and a well-behaved app writes almost nothing, so a dropped binary and a new crontab stand out:

ubuntu@secopslog-docker:~/lab/forensics · Docker 29.8.2
$ docker diff lab-suspect | tee ir-42/diff.txt
C /etc A /etc/crontab C /tmp A /tmp/.miner

A is added, C is changed. Take the whole filesystem as a tarball for offline analysis, and commit a runnable snapshot you can detonate later in an isolated lab. docker save writes that snapshot to a file, so the evidence directory holds it too and does not depend on the image staying on this host:

ubuntu@secopslog-docker:~/lab/forensics · Docker 29.8.2
$ docker export lab-suspect -o ir-42/rootfs.tar tar tf ir-42/rootfs.tar | grep -E "tmp/\.miner|etc/crontab"
etc/crontab etc/crontabs/ etc/crontabs/root tmp/.miner
$ docker commit lab-suspect lab-evidence:incident-42 docker save lab-evidence:incident-42 -o ir-42/image.tar ls ir-42
sha256:876ab58b86439f29988e62348e30715fb8acc82bd11ce21f0f467f9ea6a4e074 diff.txt image.tar logs.txt pid.txt sockets.txt environ.txt inspect.json miner.bin rootfs.tar

The export holds the dropped binary and the new crontab. Both the export and the saved image are for a lab, never for re-running attacker code on a production host. Last, hash everything you took. A copied file is a copy; a copied file with a checksum recorded at collection time is something an auditor will still trust months later:

ubuntu@secopslog-docker:~/lab/forensics · Docker 29.8.2
$ (cd ir-42 && sha256sum * > SHA256SUMS) awk '{print substr($1,1,16), $2}' ir-42/SHA256SUMS
ae9d32ae95556242 diff.txt e365c9a153ee6442 environ.txt 15396e37325d87b5 image.tar 4f88c22caf75f857 inspect.json 9246097ee5394117 logs.txt da988cc7d4ab9d83 miner.bin 4d6660a94c867ad7 pid.txt 2b31e82f8aad1a3c rootfs.tar d0b5990651786e08 sockets.txt

The subshell keeps your shell in ~/lab/forensics. The hashes vary from run to run (timestamps, PIDs and addresses differ); what matters is that SHA256SUMS is written now and travels with the files.

Memory needs CRIU, and the lab does not install it

Disk shows what was written down; a payload that lives only in RAM leaves nothing there. Capturing it means checkpointing the running process tree with CRIU, which docker checkpoint drives:

docker checkpoint (reference)
# experimental daemon + CRIU installed on the host
$ docker checkpoint create --leave-running suspicious cp-incident-42
$ docker checkpoint ls suspicious

docker checkpoint is experimental on Docker 29: the daemon must run in experimental mode, and CRIU must be installed on the host. CRIU supports x86_64, arm64 and other architectures; the lab VM simply does not install it, so this course does not run it. --checkpoint-dir chooses where the images land; without it they go under the container's directory in the daemon's data root, which docker rm deletes, so set it to your evidence storage. Without CRIU you stop at the filesystem, configuration, log and socket evidence above, which is enough to answer most of the case.

Investigate, then remediate

Lay the pieces side by side: docker diff and the recovered binary show what ran and what it dropped; the socket table and the environment show who it called; the logs and the inspect record show how it was started and what it printed; your host auditd or Falco trail shows when it arrived. Only then remediate, and remediate never means restart. Before removing anything, neutralise the restart policy so the daemon cannot bring a copy back over your evidence:

ubuntu@secopslog-docker:~/lab/forensics · Docker 29.8.2
$ docker inspect -f "{{.HostConfig.RestartPolicy.Name}}" lab-suspect docker update --restart=no lab-suspect >/dev/null docker inspect -f "{{.HostConfig.RestartPolicy.Name}}" lab-suspect
unless-stopped no
Watch out
A restart policy, or an orchestrator, exists to replace a failed container so desired state returns, which is the one thing you cannot allow mid-incident. If the suspect process dies by itself, or someone kills its host PID, a restart=always or unless-stopped policy re-runs the entrypoint: live memory is gone and the attacker's persistence may fire on the way up. (docker stop and docker kill count as a manual stop and do not trigger the policy.) Under Swarm or Kubernetes the control plane replaces the task and garbage-collects the old one, writable layer and all, while you are still capturing. Take the workload out of reconciliation without deleting it. On Kubernetes, relabel the pod so its ReplicaSet no longer selects it: the controller starts a replacement and leaves the suspect pod running, and its Service stops sending it traffic. Then isolate it with a NetworkPolicy and cordon the node. Never scale the controller to zero, which deletes the pod. For a plain container, docker update --restart=no, as above.

Reconstruct the timeline from the daemon's own event stream, then remove the container. The evidence directory is what survives:

ubuntu@secopslog-docker:~/lab/forensics · Docker 29.8.2
$ docker events --since 10m --until 0s --filter container=lab-suspect \ --format "{{.Time}} {{.Type}} {{.Action}}"
1791413823 container create 1791413824 container start 1791413826 container pause 1791413826 container archive-path 1791413827 container export
$ docker unpause lab-suspect >/dev/null docker rm -f lab-suspect docker inspect lab-suspect 2>&1 | tail -1 ls ir-42 | wc -l
lab-suspect error: no such object: lab-suspect 10

docker events with --since and --until gives each state change with a Unix timestamp: the create and start, the pause, the docker cp (archive-path) and the export. It did not list the update or the commit here, so your own case notes with times are part of the record. Stream events to your log pipeline in production, because the daemon keeps only a short history and loses it on restart. After docker rm the container, its log file and its writable layer are gone, and the ten files in ir-42 are what is left. With the evidence off the box, patch the image or close the way in, rotate every credential the container could reach and treat them as burned, redeploy from a freshly built clean image, and turn the drop you found (a new binary under /tmp writing to /etc/crontab) into a Falco rule so the next attempt trips an alarm. Clean up:

ubuntu@secopslog-docker:~/lab/forensics · Docker 29.8.2
$ docker rm -f lab-c2 >/dev/null docker rmi lab-evidence:incident-42 >/dev/null docker network rm lab-evnet rm -rf ir-42
lab-evnet
Quick check
01You have frozen a compromised container. Why read its binary and open sockets from the host (/proc/<pid>/exe, nsenter -n ss) rather than running ps, ss and ls inside it?
Correct — Whoever owns the container owns its userspace. The host's tools reading kernel-maintained entries give you facts the attacker cannot edit; cmdline and environ are the exception, since they come from process memory.
Incorrect — The issue is trust, not speed; fast answers from a compromised toolset are still wrong.
Incorrect — Reading another process's /proc entries takes host root, and exec can run as any user.
Incorrect — Pause does block exec, but that is a side effect; even on a running container you would not trust its own binaries.
02In what order should the first containment steps run, given a suspect container with one established outbound connection?
Incorrect — A paused container refuses exec, the in-container ss is untrusted, and cutting first destroys the connection the evidence is about.
Incorrect — docker rm deletes the writable layer, the json-file log and the inspect record; there is nothing left to compare.
Incorrect — Stopping kills the process, so memory and the socket table are gone before you have recorded them.
Correct — Freezing stops the process without killing it, the records disappear with docker rm, and the socket table must be saved before the interface goes.
03A compromised container runs with --restart always. Someone runs kill -9 on its host PID on reflex. What happens to the investigation?
Incorrect — A restart runs the entrypoint again in a new process; the old memory and sockets are gone.
Correct — A process that dies on its own (or is killed from the host) triggers the policy; only the writable layer survives until docker rm.
Incorrect — Restart policies restore the desired state; they never preserve a failed process.
Incorrect — It is the reverse: memory dies with the process, while the writable layer stays until the container is removed.

Try this

Work through “Investigate, then remediate” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from container forensics and incident response, keep “Investigate, then remediate”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related