Live triage: contain without destroying evidence

Order of volatility and network quarantine.

Advanced16 min · lesson 13 of 15

When a detection fires on a live host, the expensive mistake is to do the fast thing: kill the process, reboot, reimage, close the ticket. That evicts the attacker from one machine and destroys the record of how they got in, what they touched, and whether they are on your other hosts. This lesson walks the first hour on a compromised Ubuntu host: capture the state a reboot would erase, in order of volatility, to an evidence directory with hashes; recover a binary the attacker deleted while it ran; cut the host off with one atomic firewall change that keeps your own access and log shipping open and undoes itself if it locks you out; and check where the logs live and what memory acquisition the host allows. Everything runs against a harmless stand-in, and the firewall step runs inside an isolated network namespace, so the machine's real network is never touched.

Order of volatility: grab what evaporates first

Evidence has a shelf life. A file on disk survives a reboot; a network connection or a process that exists only in memory is gone in minutes and certainly gone when the power blinks. RFC 3227, the standing guideline for evidence collection, calls this the order of volatility: collect the fastest-fading things first, so the process table and sockets (seconds), then memory (minutes to hours), then the disk. It adds two rules people break under pressure: do not shut down before collecting, because shutdown scripts may be rigged to destroy evidence, and do not trust the programs on a possibly-subverted host. Start by writing down the time in a format nobody can argue about, then snapshot the perishable state into a root-only evidence directory.

deploy@web01 · Ubuntu 26.04 LTS
$ EVID=/var/tmp/det-triage/evid sudo install -d -m 700 "$EVID" sudo sh -c "date -u +%FT%TZ > $EVID/00-start-utc.txt ps -efww > $EVID/ps.txt ss -tuwanp > $EVID/sockets.txt loginctl list-sessions > $EVID/sessions.txt cat /proc/modules > $EVID/modules.txt cat /proc/sys/kernel/tainted > $EVID/tainted.txt nft list ruleset > $EVID/nft-ruleset.txt" sudo ls -lh /var/tmp/det-triage/evid
total 36K -rw-r--r-- 1 root root 21 Sep 27 08:40 00-start-utc.txt -rw-r--r-- 1 root root 3.4K Sep 27 08:40 modules.txt -rw-r--r-- 1 root root 277 Sep 27 08:40 nft-ruleset.txt -rw-r--r-- 1 root root 12K Sep 27 08:40 ps.txt -rw-r--r-- 1 root root 186 Sep 27 08:40 sessions.txt -rw-r--r-- 1 root root 2.2K Sep 27 08:40 sockets.txt -rw-r--r-- 1 root root 2 Sep 27 08:40 tainted.txt

Seven files in seconds: the start time, every process with its arguments, every socket (-tuwanp covers TCP, UDP, raw and Unix sockets, where -tanp would show TCP only), the login sessions, the loaded kernel modules, the kernel's taint value, and the firewall ruleset as it stood before you added anything to it, which preserves any rule the attacker added. On a real incident this directory is on external media or streamed to a collector, because writing to the disk you are investigating changes it. The module list is only useful if you know what a suspicious line looks like.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo cat /var/tmp/det-triage/evid/tainted.txt sudo grep -E "\([A-Z]+\)$" /var/tmp/det-triage/evid/modules.txt || echo "no module carries a taint flag"
0 no module carries a taint flag

0 means the kernel is untainted, and no module line carries a flag. A non-zero value, or a module line ending in flags such as (O) for an out-of-tree module or (E) for an unsigned one, deserves a closer look on a host that should run only distribution modules (linux-hard/modules explains the letters). One more item needs a check of its own: where the list of logged-in users comes from.

deploy@web01 · Ubuntu 26.04 LTS
$ echo "who (utmp): $(who | wc -l) sessions" echo "loginctl (logind): $(loginctl list-sessions --no-legend | wc -l) sessions"
who (utmp): 0 sessions loginctl (logind): 2 sessions

who reads the utmp database, and on Ubuntu 26.04 it reports nothing even though systemd-logind tracks two live sessions. A collection script written for older systems that saves who -a would record an empty user list, which is why the capture asks loginctl list-sessions. Check what your platform populates before an incident, not during one.

Follow the connection to a deleted binary

Malware that phones home has to open a socket, and every socket has a process behind it. A common tell is a process running from a file that no longer exists: the payload copied itself into a scratch directory, started, then deleted its own file. The kernel still holds a link to the original bytes at /proc/<pid>/exe and marks it deleted. The lab started a stand-in this way, a copy of gnusleep named .kworkerd (the Try this section shows the three commands).

deploy@web01 · Ubuntu 26.04 LTS
$ pid=$(pgrep -u deploy -f "det-triage/run/.kworkerd" | head -1) sudo ls -l /proc/$pid/exe
lrwxrwxrwx 1 deploy deploy 0 Sep 27 08:40 /proc/85793/exe -> /var/tmp/det-triage/run/.kworkerd (deleted)

A hidden name in a scratch directory, and (deleted) because the file was removed after the process started. The bytes are still mapped, so copy the program straight out of /proc and hash it before the process exits; it is often the only chance you get at the sample.

deploy@web01 · Ubuntu 26.04 LTS
$ pid=$(pgrep -u deploy -f "det-triage/run/.kworkerd" | head -1) sudo cp /proc/$pid/exe /var/tmp/det-triage/evid/pid-$pid-kworkerd.elf sudo sha256sum /var/tmp/det-triage/evid/pid-$pid-kworkerd.elf
ccadf05b109841bf88ae2be395714f45be62f719e436b15e2b0a11aaceec0b88 /var/tmp/det-triage/evid/pid-85793-kworkerd.elf

Take the rest of the process's own state while you are there: its full command line, environment, status, memory map, working directory and open file descriptors, all readable from /proc/<pid>/.

deploy@web01 · Ubuntu 26.04 LTS
$ pid=$(pgrep -u deploy -f "det-triage/run/.kworkerd" | head -1) EVID=/var/tmp/det-triage/evid for f in cmdline environ status maps; do sudo cat /proc/$pid/$f | sudo tee $EVID/pid-$pid-$f >/dev/null; done sudo ls -l /proc/$pid/cwd /proc/$pid/fd/ | sudo tee $EVID/pid-$pid-fd.txt sudo cat $EVID/pid-$pid-cmdline | tr "\0" " "; echo
lrwxrwxrwx 1 deploy deploy 0 Sep 27 08:40 /proc/85793/cwd -> /home/deploy /proc/85793/fd/: total 0 lr-x------ 1 deploy deploy 64 Sep 27 08:40 0 -> /dev/null l-wx------ 1 deploy deploy 64 Sep 27 08:40 1 -> /dev/null l-wx------ 1 deploy deploy 64 Sep 27 08:40 2 -> /dev/null /var/tmp/det-triage/run/.kworkerd 900

Its descriptors point at /dev/null and its working directory is deploy's home; a real implant would show sockets and files here that name its peers and its staging area. For the process's memory itself, gcore from the gdb package writes a core file of one process, which is the practical fallback on hosts where full-memory tools are blocked (below). None of this works once the process is gone. If it must stop acting before containment is in place, freeze it rather than kill it: sudo kill -STOP <pid>, or sudo systemctl freeze <unit> for a service, keeps its memory and sockets intact. Some samples watch for being stopped, so freezing is a trade-off too.

Isolate without pulling the plug

With the seconds-long state captured, contain. If the host is actively harming others (exfiltrating, scanning, attacking), do it now and image memory afterwards: a full image of a large host takes minutes to hours, and isolation above the host (a quarantine VLAN or a locked-down security group) destroys no volatile state. A host firewall is the fast first move. Its advantage over pulling the cable is what it keeps open: your collector and admin session, and the log shipping the rest of this course builds (audisp-remote, the log shipper, osquery and Falco outputs, an EDR agent). To the malware, both look alike, since its channel simply stops answering.

The lockout risk is real. Before loading, read your own session's address as the host sees it (echo $SSH_CONNECTION, or ss -tnp state established '( sport = :22 )'): behind a NAT, bastion or VPN it is not your laptop's address. Have console or serial access ready, and arm a rollback. The demonstration runs in a network namespace, an isolated network stack, built with these commands; 198.51.100.1 stands in for the collector and your admin session, 198.51.100.3 for the log server (a listener on TCP 6514, syslog over TLS), and 198.51.100.9 for the attacker's command-and-control (C2) address.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns add det-tri-ns sudo ip link add det-tri-h type veth peer name det-tri-n sudo ip link set det-tri-n netns det-tri-ns sudo ip addr add 198.51.100.1/24 dev det-tri-h sudo ip addr add 198.51.100.3/24 dev det-tri-h sudo ip addr add 198.51.100.9/24 dev det-tri-h sudo ip link set det-tri-h up sudo ip netns exec det-tri-ns ip addr add 198.51.100.2/24 dev det-tri-n sudo ip netns exec det-tri-ns ip link set det-tri-n up sudo ip netns exec det-tri-ns ip link set lo up sudo setsid -f ncat -lk --recv-only 198.51.100.3 6514 </dev/null >/dev/null 2>&1
$ sudo ip netns exec det-tri-ns ping -c1 -W1 198.51.100.1 | tail -2 sudo ip netns exec det-tri-ns nc -zv -w2 198.51.100.3 6514 sudo ip netns exec det-tri-ns ping -c1 -W1 198.51.100.9 | tail -2
1 packets transmitted, 1 received, 0% packet loss, time 0ms rtt min/avg/max/mdev = 0.017/0.017/0.017/0.000 ms Connection to 198.51.100.3 6514 port [tcp/syslog-tls] succeeded! 1 packets transmitted, 1 received, 0% packet loss, time 0ms rtt min/avg/max/mdev = 0.045/0.045/0.045/0.000 ms

All three answer. Arm the rollback first: systemd-run creates a transient timer that deletes the quarantine table in five minutes unless you cancel it, so a rule that locks you out undoes itself.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --unit=det-tri-undo --on-active=5min ip netns exec det-tri-ns nft delete table inet quar
Running timer as unit: det-tri-undo.timer Will run service as unit: det-tri-undo.service

The quarantine is one file, loaded in one nftables transaction.

/var/tmp/det-triage/quar.nft
table inet quar {
chain input {
type filter hook input priority -400; policy drop;
iif "lo" accept
ip saddr 198.51.100.1 accept
ip saddr 198.51.100.3 tcp sport 6514 accept
}
chain output {
type filter hook output priority -400; policy drop;
oif "lo" accept
ip daddr 198.51.100.1 accept
ip daddr 198.51.100.3 tcp dport 6514 accept
}
}

Both chains drop anything not explicitly allowed. The collector gets full access; the log server only its port. priority -400 places the table ahead of ordinary filter chains, because in nftables a drop is final while an accept is not: a later chain could still accept traffic you meant to block, so you run first and drop hard. Because nft -f commits the whole file at once, there is no instant where the drop is live and the exceptions are not, and nft -c checks the file without loading it. Add your time server the same way if the triage will run long.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec det-tri-ns nft -c -f /var/tmp/det-triage/quar.nft sudo ip netns exec det-tri-ns nft -f /var/tmp/det-triage/quar.nft sudo ip netns exec det-tri-ns nft list tables
table inet quar
$ sudo ip netns exec det-tri-ns ping -c1 -W1 198.51.100.1 | tail -2 sudo ip netns exec det-tri-ns nc -zv -w2 198.51.100.3 6514
1 packets transmitted, 1 received, 0% packet loss, time 0ms rtt min/avg/max/mdev = 0.021/0.021/0.021/0.000 ms Connection to 198.51.100.3 6514 port [tcp/syslog-tls] succeeded!
$ sudo ip netns exec det-tri-ns ping -c1 -W2 198.51.100.9
PING 198.51.100.9 (198.51.100.9) 56(84) bytes of data. --- 198.51.100.9 ping statistics --- 1 packets transmitted, 0 received, 100% packet loss, time 0ms

The collector still answers, the log port still connects, and the C2 address is cut. Access is confirmed, so cancel the rollback; no timer remains.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop det-tri-undo.timer systemctl list-timers --all --no-legend det-tri-undo.timer | wc -l
0

A rollback you have never seen fire is a guess, so rehearse it on a test host: arm it with a short delay and let it run.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --unit=det-tri-undo --on-active=20s ip netns exec det-tri-ns nft delete table inet quar
Running timer as unit: det-tri-undo.timer Will run service as unit: det-tri-undo.service
$ journalctl -u det-tri-undo.service -o cat --no-pager | tail -2 sudo ip netns exec det-tri-ns nft list tables | wc -l sudo ip netns exec det-tri-ns ping -c1 -W1 198.51.100.9 | tail -2
Started det-tri-undo.service - [systemd-run] /usr/sbin/ip netns exec det-tri-ns nft delete table inet quar. det-tri-undo.service: Deactivated successfully. 0 1 packets transmitted, 1 received, 0% packet loss, time 0ms rtt min/avg/max/mdev = 0.256/0.256/0.256/0.000 ms

The journal shows the transient service ran the delete, the namespace has no tables left, and the C2 address answers again. By default systemd may run a timer up to a minute late (AccuracySec=, systemd.timer(5)), so allow for that when you pick the delay. Treat the host firewall as a first move only: anyone with root on the box, possibly the attacker, can delete the table with nft. Authoritative isolation belongs above the host, where the intruder's root cannot reach it.

Where the logs live

"Do not reboot" becomes a hard rule on hosts whose journal lives in RAM. journald's Storage= setting decides: persistent writes to /var/log/journal and creates it if needed, volatile keeps the journal in /run/log/journal, a tmpfs lost on reboot, and auto persists only if /var/log/journal already exists. The default is chosen when systemd is built, so it differs by distribution. Check the effective setting, and where the active files are.

deploy@web01 · Ubuntu 26.04 LTS
$ systemd-analyze cat-config systemd/journald.conf | grep -E "^#?Storage=" ls -A /run/log/journal /var/log/journal journalctl --disk-usage
#Storage=persistent /run/log/journal: /var/log/journal: 7fa05bc53e6b40f880f494631079f2df Archived and active journals take up 125.3M in the file system.
deploy@rocky10 · Rocky Linux 10.2
$ systemd-analyze cat-config systemd/journald.conf | grep -E "^#?Storage=" rpm -qf /var/log/journal
#Storage=auto systemd-257-23.el10_2.2.rocky.0.1.aarch64

Ubuntu 26.04's systemd 259 is built with persistent, the commented line in the shipped journald.conf shows it, and /run/log/journal is empty because the active files are on disk. Rocky Linux 10.2's systemd 257 is built with auto, and its systemd package ships /var/log/journal, so it persists too (journald.conf(5) on each VM). Images that remove the directory or set volatile, as some minimal and container images do, lose everything on reboot. Either way, copy the logs to your evidence.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo sh -c "journalctl --no-pager -o export --since \"-1h\" > /var/tmp/det-triage/evid/journal.export cp -p /var/log/auth.log /var/tmp/det-triage/evid/" sudo ls -lh /var/tmp/det-triage/evid/journal.export /var/tmp/det-triage/evid/auth.log
-rw-r----- 1 syslog adm 5.3M Sep 27 08:41 /var/tmp/det-triage/evid/auth.log -rw-r--r-- 1 root root 16M Sep 27 08:41 /var/tmp/det-triage/evid/journal.export

-o export keeps every field for later tools. The lab exports only the last hour to stay small; on a real incident export from well before the first known activity, or copy the journal files themselves, and take /var/log/audit/ and auth.log (secure on RHEL) as well.

Memory is a plan, not a reflex

Memory holds what never touches the disk: keys, decrypted payloads, injected code, the real arguments of a program that rewrote its command line. Acquiring it is not a one-liner you can assume will work, because the hardening that stops kernel rootkits also constrains the tools that read memory. Read the host's constraints first.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /sys/kernel/security/lockdown cat /sys/module/module/parameters/sig_enforce uname -m
[none] integrity confidentiality N aarch64
Full-memory tools and what blocks them
AVML (Microsoft)
static userspace binary
reads /dev/crash, /proc/kcore or /dev/mem
x86_64 only
will not run on this aarch64 host
blocked by lockdown
per its README
LEMON (EURECOM)
eBPF-based acquirer
no kernel module needed
x86_64 and ARM64
blocked by confidentiality
works under integrity lockdown
LiME
loadable kernel module
inserted with insmod
refused if unsigned
under sig_enforce or lockdown
blocked by modules_disabled
unless built and signed in advance
Pick the tool from the host's architecture, lockdown mode and module policy, and test it before you need it.

This host is aarch64 with lockdown [none] and sig_enforce N, so AVML is out, while LEMON or an unsigned LiME would work. On a production host hardened the way linux-hard/modules recommends, only LEMON, a LiME build signed with an enrolled key in advance, or a snapshot from the hypervisor below the guest would. Decide which before the incident, and test it on a matching host.

Your own actions are evidence too

Chain of custody is the paper trail that lets someone else trust your evidence: who collected what, when, and that it has not changed since. When collection is done, freeze a manifest, a SHA-256 sum of every artefact, stored beside the evidence.

deploy@web01 · Ubuntu 26.04 LTS
$ EVID=/var/tmp/det-triage/evid sudo sh -c "cd $EVID && find . -maxdepth 1 -type f ! -name manifest.sha256 -exec sha256sum {} + | sort -k2 > manifest.sha256" sudo wc -l /var/tmp/det-triage/evid/manifest.sha256 sudo cut -c1-20 /var/tmp/det-triage/evid/manifest.sha256
15 /var/tmp/det-triage/evid/manifest.sha256 aa31e98ab607b61a6db8 c3d939ef883130574e6e 1f7110c583da995ae7e0 2f62f02fb697f553d4e4 7a1fa2dd4ca3e0204380 287e0aa9214c18ae351c 913fe9e9b0d1ce30a5a6 e0f22c79f4ab4426caa6 ccadf05b109841bf88ae 0474f5a1169b67b28b73 fa8a2ee50fc8adbe5b82 9c40febf71faadf2c988 705aa1ff010ed153d081 199b38fd04fed4b1b15f 9a271f2a916b0b6ee6ce

Fifteen artefacts, fifteen sums (the preview shows the first 20 hex characters of each). The ninth, ccadf05b..., is the hash you took when recovering the binary from /proc. Anyone can recompute the set later with sha256sum -c manifest.sha256 and prove nothing changed after collection. Send a copy of the manifest somewhere the suspect host cannot write.

The first hour, in order
1Record the time
UTC, before the first real command
2Snapshot live state
processes, sockets, modules, users, rules
3Recover the process
binary and /proc files, freeze if needed
4Isolate
upstream first; host firewall with a rollback
5Preserve logs, then memory
journal and logs off the host, then RAM
6Hash and hand off
a manifest, then begin analysis
Seconds-long captures come first, containment next, long acquisitions after; reboot last, if at all.

Try this

On a lab host, practise the capture-before-you-touch discipline. Make an evidence directory with sudo install -d -m 700, write date -u +%FT%TZ into it, then save ps -efww, ss -tuwanp, loginctl list-sessions and cat /proc/modules to files there. Start a harmless background process from a scratch file and delete the file: cp /usr/bin/gnusleep /var/tmp/x; setsid -f /var/tmp/x 600; rm /var/tmp/x (on Ubuntu 26.04 use gnusleep, because the default sleep is part of the uutils multi-call binary and will not run under another name). Get its PID with pgrep -f /var/tmp/x, confirm ls -l /proc/<pid>/exe reads (deleted), and copy the bytes back out with cp /proc/<pid>/exe recovered.bin. Then build the namespace from the lesson, arm the rollback with a one-minute delay, load quar.nft, confirm the collector answers and the C2 address does not, and watch the table disappear when the timer fires. Finish with sha256sum of every file into manifest.sha256, check it with sha256sum -c manifest.sha256, and remove the namespace with sudo ip netns del det-tri-ns.

Takeaway

Capture the seconds-long state first (the process table, sockets and the running binary from /proc), and do not let a long memory image delay containment of a host that is harming others: isolate upstream, which destroys nothing, then acquire. On the host itself, quarantine with one atomic firewall load that keeps your collector and log shipping reachable, arm a rollback before you load it, and reboot last, if at all.

Quick check
01A process calling itself kworkerd is beaconing to a foreign address. Its /proc/<pid>/exe points at a deleted file in a scratch directory, and the host keeps its journal only in /run/log/journal. What do you do first?
Incorrect — The cable takes your collector and log shipping with it, so evidence can only come off by hand, and nothing captured the process before the plug came out.
Incorrect — A journal in /run/log/journal lives in tmpfs, so the reboot that gives you a clean host also erases the very history you planned to read.
Incorrect — The instant it dies the /proc link dies with it, and the file was already unlinked, so there is nothing left on disk to recover, and some samples wipe themselves on termination.
Correct — The seconds-long capture comes first because the link lives only as long as the process; upstream isolation then stops the beacon without destroying memory, and the long image runs after that.
02You are quarantining a live host with nftables while pulling evidence over SSH from your collector. A colleague suggests running the drop-everything policy as one nft command and the collector-allow rule as a second command right after. What is wrong with that?
Incorrect — Each nft command is its own transaction and commits the moment it returns, so the first commit is exactly the window that cuts you off.
Correct — One nft -f commits the drop policy and the collector exception together, so there is never a moment where everything is dropped and the collector has not yet been permitted.
Incorrect — Adding a rule does not delete an earlier one; the hazard is the timing between two separate commits, not one command cancelling another.
Incorrect — Both rules go into the same table at the same priority; the real problem is the window between two commits, not a priority mismatch.
03Your incident host runs signed kernel modules with sig_enforce=Y and kernel lockdown in integrity mode. You planned to acquire memory with an unsigned LiME build. What does this posture mean for that plan?
Incorrect — Signature enforcement applies to every module load, boot-time or manual; an unsigned LiME will be refused exactly when you try to insert it.
Incorrect — LiME is a loadable kernel module, not a userspace tool; that is precisely why module-signing and lockdown constrain it.
Correct — The hardening that blocks a kernel rootkit blocks an unsigned module, so acquisition is designed around it: a build signed with an enrolled key, an eBPF acquirer that works in integrity mode, or a capture from below the guest.
Incorrect — Integrity mode blocks modifying the running kernel and unsigned modules; it is confidentiality mode that also blocks reads like /proc/kcore, and eBPF or hypervisor capture can still work.

Related