Live triage: contain without destroying evidence
Order of volatility and network quarantine.
When a detection fires on a live host, the expensive mistake is to do the fast thing: kill the process, reboot, reimage, close the ticket. That evicts the attacker from one machine and destroys the record of how they got in, what they touched, and whether they are on your other hosts. This lesson walks the first hour on a compromised Ubuntu host: capture the state a reboot would erase, in order of volatility, to an evidence directory with hashes; recover a binary the attacker deleted while it ran; cut the host off with one atomic firewall change that keeps your own access and log shipping open and undoes itself if it locks you out; and check where the logs live and what memory acquisition the host allows. Everything runs against a harmless stand-in, and the firewall step runs inside an isolated network namespace, so the machine's real network is never touched.
Order of volatility: grab what evaporates first
Evidence has a shelf life. A file on disk survives a reboot; a network connection or a process that exists only in memory is gone in minutes and certainly gone when the power blinks. RFC 3227, the standing guideline for evidence collection, calls this the order of volatility: collect the fastest-fading things first, so the process table and sockets (seconds), then memory (minutes to hours), then the disk. It adds two rules people break under pressure: do not shut down before collecting, because shutdown scripts may be rigged to destroy evidence, and do not trust the programs on a possibly-subverted host. Start by writing down the time in a format nobody can argue about, then snapshot the perishable state into a root-only evidence directory.
Seven files in seconds: the start time, every process with its arguments, every socket (-tuwanp covers TCP, UDP, raw and Unix sockets, where -tanp would show TCP only), the login sessions, the loaded kernel modules, the kernel's taint value, and the firewall ruleset as it stood before you added anything to it, which preserves any rule the attacker added. On a real incident this directory is on external media or streamed to a collector, because writing to the disk you are investigating changes it. The module list is only useful if you know what a suspicious line looks like.
0 means the kernel is untainted, and no module line carries a flag. A non-zero value, or a module line ending in flags such as (O) for an out-of-tree module or (E) for an unsigned one, deserves a closer look on a host that should run only distribution modules (linux-hard/modules explains the letters). One more item needs a check of its own: where the list of logged-in users comes from.
who reads the utmp database, and on Ubuntu 26.04 it reports nothing even though systemd-logind tracks two live sessions. A collection script written for older systems that saves who -a would record an empty user list, which is why the capture asks loginctl list-sessions. Check what your platform populates before an incident, not during one.
Follow the connection to a deleted binary
Malware that phones home has to open a socket, and every socket has a process behind it. A common tell is a process running from a file that no longer exists: the payload copied itself into a scratch directory, started, then deleted its own file. The kernel still holds a link to the original bytes at /proc/<pid>/exe and marks it deleted. The lab started a stand-in this way, a copy of gnusleep named .kworkerd (the Try this section shows the three commands).
A hidden name in a scratch directory, and (deleted) because the file was removed after the process started. The bytes are still mapped, so copy the program straight out of /proc and hash it before the process exits; it is often the only chance you get at the sample.
Take the rest of the process's own state while you are there: its full command line, environment, status, memory map, working directory and open file descriptors, all readable from /proc/<pid>/.
Its descriptors point at /dev/null and its working directory is deploy's home; a real implant would show sockets and files here that name its peers and its staging area. For the process's memory itself, gcore from the gdb package writes a core file of one process, which is the practical fallback on hosts where full-memory tools are blocked (below). None of this works once the process is gone. If it must stop acting before containment is in place, freeze it rather than kill it: sudo kill -STOP <pid>, or sudo systemctl freeze <unit> for a service, keeps its memory and sockets intact. Some samples watch for being stopped, so freezing is a trade-off too.
Isolate without pulling the plug
With the seconds-long state captured, contain. If the host is actively harming others (exfiltrating, scanning, attacking), do it now and image memory afterwards: a full image of a large host takes minutes to hours, and isolation above the host (a quarantine VLAN or a locked-down security group) destroys no volatile state. A host firewall is the fast first move. Its advantage over pulling the cable is what it keeps open: your collector and admin session, and the log shipping the rest of this course builds (audisp-remote, the log shipper, osquery and Falco outputs, an EDR agent). To the malware, both look alike, since its channel simply stops answering.
The lockout risk is real. Before loading, read your own session's address as the host sees it (echo $SSH_CONNECTION, or ss -tnp state established '( sport = :22 )'): behind a NAT, bastion or VPN it is not your laptop's address. Have console or serial access ready, and arm a rollback. The demonstration runs in a network namespace, an isolated network stack, built with these commands; 198.51.100.1 stands in for the collector and your admin session, 198.51.100.3 for the log server (a listener on TCP 6514, syslog over TLS), and 198.51.100.9 for the attacker's command-and-control (C2) address.
All three answer. Arm the rollback first: systemd-run creates a transient timer that deletes the quarantine table in five minutes unless you cancel it, so a rule that locks you out undoes itself.
The quarantine is one file, loaded in one nftables transaction.
table inet quar {chain input {type filter hook input priority -400; policy drop;iif "lo" acceptip saddr 198.51.100.1 acceptip saddr 198.51.100.3 tcp sport 6514 accept}chain output {type filter hook output priority -400; policy drop;oif "lo" acceptip daddr 198.51.100.1 acceptip daddr 198.51.100.3 tcp dport 6514 accept}}
Both chains drop anything not explicitly allowed. The collector gets full access; the log server only its port. priority -400 places the table ahead of ordinary filter chains, because in nftables a drop is final while an accept is not: a later chain could still accept traffic you meant to block, so you run first and drop hard. Because nft -f commits the whole file at once, there is no instant where the drop is live and the exceptions are not, and nft -c checks the file without loading it. Add your time server the same way if the triage will run long.
The collector still answers, the log port still connects, and the C2 address is cut. Access is confirmed, so cancel the rollback; no timer remains.
A rollback you have never seen fire is a guess, so rehearse it on a test host: arm it with a short delay and let it run.
The journal shows the transient service ran the delete, the namespace has no tables left, and the C2 address answers again. By default systemd may run a timer up to a minute late (AccuracySec=, systemd.timer(5)), so allow for that when you pick the delay. Treat the host firewall as a first move only: anyone with root on the box, possibly the attacker, can delete the table with nft. Authoritative isolation belongs above the host, where the intruder's root cannot reach it.
Where the logs live
"Do not reboot" becomes a hard rule on hosts whose journal lives in RAM. journald's Storage= setting decides: persistent writes to /var/log/journal and creates it if needed, volatile keeps the journal in /run/log/journal, a tmpfs lost on reboot, and auto persists only if /var/log/journal already exists. The default is chosen when systemd is built, so it differs by distribution. Check the effective setting, and where the active files are.
Ubuntu 26.04's systemd 259 is built with persistent, the commented line in the shipped journald.conf shows it, and /run/log/journal is empty because the active files are on disk. Rocky Linux 10.2's systemd 257 is built with auto, and its systemd package ships /var/log/journal, so it persists too (journald.conf(5) on each VM). Images that remove the directory or set volatile, as some minimal and container images do, lose everything on reboot. Either way, copy the logs to your evidence.
-o export keeps every field for later tools. The lab exports only the last hour to stay small; on a real incident export from well before the first known activity, or copy the journal files themselves, and take /var/log/audit/ and auth.log (secure on RHEL) as well.
Memory is a plan, not a reflex
Memory holds what never touches the disk: keys, decrypted payloads, injected code, the real arguments of a program that rewrote its command line. Acquiring it is not a one-liner you can assume will work, because the hardening that stops kernel rootkits also constrains the tools that read memory. Read the host's constraints first.
This host is aarch64 with lockdown [none] and sig_enforce N, so AVML is out, while LEMON or an unsigned LiME would work. On a production host hardened the way linux-hard/modules recommends, only LEMON, a LiME build signed with an enrolled key in advance, or a snapshot from the hypervisor below the guest would. Decide which before the incident, and test it on a matching host.
Your own actions are evidence too
Chain of custody is the paper trail that lets someone else trust your evidence: who collected what, when, and that it has not changed since. When collection is done, freeze a manifest, a SHA-256 sum of every artefact, stored beside the evidence.
Fifteen artefacts, fifteen sums (the preview shows the first 20 hex characters of each). The ninth, ccadf05b..., is the hash you took when recovering the binary from /proc. Anyone can recompute the set later with sha256sum -c manifest.sha256 and prove nothing changed after collection. Send a copy of the manifest somewhere the suspect host cannot write.
Try this
On a lab host, practise the capture-before-you-touch discipline. Make an evidence directory with sudo install -d -m 700, write date -u +%FT%TZ into it, then save ps -efww, ss -tuwanp, loginctl list-sessions and cat /proc/modules to files there. Start a harmless background process from a scratch file and delete the file: cp /usr/bin/gnusleep /var/tmp/x; setsid -f /var/tmp/x 600; rm /var/tmp/x (on Ubuntu 26.04 use gnusleep, because the default sleep is part of the uutils multi-call binary and will not run under another name). Get its PID with pgrep -f /var/tmp/x, confirm ls -l /proc/<pid>/exe reads (deleted), and copy the bytes back out with cp /proc/<pid>/exe recovered.bin. Then build the namespace from the lesson, arm the rollback with a one-minute delay, load quar.nft, confirm the collector answers and the C2 address does not, and watch the table disappear when the timer fires. Finish with sha256sum of every file into manifest.sha256, check it with sha256sum -c manifest.sha256, and remove the namespace with sudo ip netns del det-tri-ns.
Takeaway
Capture the seconds-long state first (the process table, sockets and the running binary from /proc), and do not let a long memory image delay containment of a host that is harming others: isolate upstream, which destroys nothing, then acquire. On the host itself, quarantine with one atomic firewall load that keeps your collector and log shipping reachable, arm a rollback before you load it, and reboot last, if at all.
kworkerd is beaconing to a foreign address. Its /proc/<pid>/exe points at a deleted file in a scratch directory, and the host keeps its journal only in /run/log/journal. What do you do first?/run/log/journal lives in tmpfs, so the reboot that gives you a clean host also erases the very history you planned to read./proc link dies with it, and the file was already unlinked, so there is nothing left on disk to recover, and some samples wipe themselves on termination.nft command and the collector-allow rule as a second command right after. What is wrong with that?nft command is its own transaction and commits the moment it returns, so the first commit is exactly the window that cuts you off.nft -f commits the drop policy and the collector exception together, so there is never a moment where everything is dropped and the collector has not yet been permitted.sig_enforce=Y and kernel lockdown in integrity mode. You planned to acquire memory with an unsigned LiME build. What does this posture mean for that plan?/proc/kcore, and eBPF or hypervisor capture can still work.