The incident response loop
Contain, eradicate, recover and learn.
This closing lesson takes one incident all the way round the response loop on a Linux host: scope it before touching anything, contain it without losing evidence, remove it together with everything that keeps it alive, prove it stays gone, and turn it into a detection you have watched fire. The incident is small and harmless, planted in the lab: an unexpected service runs a renamed copy of a network tool out of /var/tmp, restarts itself when killed, and has a watchdog that starts it again. A real implant needs the same moves; only the stakes differ.
The loop has six steps: preparation, detection and scoping, containment, eradication, recovery, and lessons learned. NIST SP 800-61 Revision 3 maps them onto the Functions of the NIST Cybersecurity Framework (CSF) 2.0, and the diagram names the CSF outcome where it helps.
Detect and scope before touching anything
The alert is the kind the earlier lessons build: a baseline diff (the hourly audit from "Finding escalation paths before an attacker does", linux-det/pedetect, reports the new service, the watchdog timer below and the hashes of their unit files) or a listener hunt (linux-det/hunting) reports a new listening socket owned by a process called det-ir-cache. The lab planted, as root, a copy of /usr/bin/ncat at /var/tmp/.det-ir/det-ir-cache, the unit shown below, and a watchdog timer and service that the scoping finds. Hold off on the reflex to kill it until you can answer four questions: what is it, when did it arrive, what keeps it running, and where else is it. Start with the unit that owns it.
The unit file sits in /etc/systemd/system, where administrators put their own units, not in /usr/lib/systemd/system where packages install theirs. It is enabled, so WantedBy=multi-user.target starts it at every boot (preset: is only the distribution's default policy and says nothing about who enabled it). The CGroup line shows the full command: a program in a hidden directory under /var/tmp, which is world-writable and survives reboots, and where no packaged service keeps its binary. Restart=always brings it back whenever the process exits. Next, the process and the file behind it.
PPID 1 on its own only says the parent is init, which is also true of an orphan that init adopted (linux-det/hunting found one). Here the PID is the unit's MainPID, which is what ps was given, so systemd started it as this service, not someone's shell. STARTED gives the time to the second. The hash is identical to /usr/bin/ncat: this is a renamed copy of the ncat network tool, which dpkg -S confirms belongs to the ncat package while no package owns the copy (rpm -qf answers the same question on RHEL). A renamed legitimate tool is common in real intrusions; the risk lies in what it is told to do, which the arguments and sockets show.
It listens on 127.0.0.1:47001 and nothing else; on a real case also check ss -tnp for established connections leaving the host. The binary and the unit file were born milliseconds apart, and reading the journal from that moment gives the unit's first Started line within a second. That is the arrival time: the anchor for reading the authentication log and the journal with the method of linux-det/artifacts, to learn who did it. The last question on this host is what else refers to it.
The name search reads the unit and cron directories, the shell startup files, and the sudoers and PAM configuration. It found det-ir-watch.service, and systemctl cat shows what it does: it runs systemctl start det-ir-cache.service, and det-ir-watch.timer runs it every 15 seconds. systemctl status never mentioned the watchdog, and TriggeredBy= is empty, because the timer triggers a different unit that merely starts this one; only a search for the name found it. The unit's properties answer the restart question without touching the process: Restart=always with RestartUSec=2s brings a killed process back in two seconds. On a fleet you ask the same questions everywhere before acting, with the hash, the path and the unit names as indicators (osquery, from linux-det/osquery, answers them across hosts), because cleaning one host while the same foothold sits on another tells the intruder exactly what you can see.
Capture first, then contain
Capture before any kill or stop, every time. The running process can hold what the disk does not (a deleted binary, decrypted configuration, open connections), a kill loses all of it, and it warns whoever operates the implant. Copy the executable out of /proc, its command line, its open files and the socket table, then every file of the persistence, into a root-only directory, and hash all of it.
The first line is the command line from /proc. The hash of pid-418823-exe, the bytes the process was actually running, matches the file on disk and /usr/bin/ncat; had the file been deleted or replaced, the /proc copy would be the only one. In a real case this directory goes off the host next. With the evidence safe, the lab shows why the two quick moves are not containment. The kill is a lab demonstration only: on a real implant the unit has already told you what it would do.
A new main PID and NRestarts=1: Restart=always replaced the process two seconds later, as RestartUSec said. The second move is stopping the unit, which Restart= does respect.
The stop held for eight seconds: the journal shows the service stopped at 13:14:11, and at 13:14:19 the timer fired, the watchdog ran systemctl start, and the service was back. Stop the watchdog timer and the service in one command, so nothing is left to revive the other.
Both units are inactive: nothing runs and nothing restarts, but everything is still installed. If the program had open connections to an outside address, you would quarantine the host right after the capture, with an nftables table that keeps your own access and log shipping open (linux-det/triage), and before any of this.
Eradicate, recover and verify
Eradication removes every piece: the enable links, the unit files, the binary, and systemd's loaded copy of the units.
disable removed the two symlinks that started the units at boot, and daemon-reload made systemd forget the deleted files. systemctl mask, which links a unit name to /dev/null so nothing can start it, is not needed here. Use it when the unit comes from a package in /usr/lib/systemd/system (a deleted packaged file returns with the next update) or when something you have not yet found may recreate it. Then prove the result, and wait longer than the watchdog's interval before you do.
Twenty seconds later neither unit exists, no file in the locations the sweep reads names the service, and nothing listens on port 47001. That proves the service is gone; it does not prove the host is clean. An intruder who could write to /etc/systemd/system had root, and no command proves a negative. Recovery for a root compromise means rebuilding from a known-good image, restoring data from before the intrusion, closing the way in, and rotating every credential the host held or that was typed on it. CSF 2.0 adds an indicator check before return to production (RC.RP-05) and an after-action report (RC.RP-06).
Where persistence lives, and how to check it
A name sweep finds only what mentions a name you already know. After a root compromise, and before you trust any rebuilt host, check every place that can make a program start again, whatever it is called. These are the usual places on Linux, grouped by what starts them.
Two references tell you whether a file there belongs. For files a package installed, the package manager: dpkg --verify (Ubuntu) or rpm -Va (RHEL) reports every packaged file whose contents changed, and dpkg -S or rpm -qf shows whether a package owns a path at all. For everything a package does not track (files generated at install time, crontabs, authorized_keys, per-user startup files and user units), a baseline taken while the host was trusted, such as the one linux-det/pedetect builds. Paths differ slightly by platform: RHEL uses /etc/bashrc, /var/spool/at and /usr/lib64/security, Ubuntu /etc/bash.bashrc, /var/spool/cron/atjobs and /usr/lib/<arch>-linux-gnu/security.
dpkg --verify checked the packages that own the cron, shell, sudo, login, PAM and SSH files, and printed one line. ??5?????? means the file's checksum no longer matches the package, and c marks a configuration file. This change is the lab VM's own: its provisioning edited /etc/pam.d/runuser-l so lab steps get a login environment. On your host, a line you cannot explain from change records is a finding. The next two lines show the limit of the package manager: /etc/profile is copied into place by an install script and common-auth is generated by pam-auth-update, so no package lists them and only a baseline can vouch for them. /etc/ld.so.preload does not exist, which is the normal state. On RHEL, rpm -Va prints the same kind of line (S.5....T. c for a changed configuration file).
The last group starts programs without any unit, job or login. A udev rule matches a device event and can run a program for it through a RUN key (udev(7)); rules in /etc/udev/rules.d are the administrator's and replace a packaged rule of the same name in /usr/lib/udev/rules.d. Every module named in a modules-load.d file is loaded at boot by systemd-modules-load.service (modules-load.d(5)), and an install line in a modprobe.d file makes modprobe run a shell command instead of loading the module it names (modprobe.d(5)). For the files a package does not track you need a baseline. This one is deliberately small: the enabled units, the loaded modules, the authorized_keys hashes, and an empty stamp file whose time marks the moment the host was trusted. The lab took it before it planted the incident, and a real one belongs off the host (linux-det/pedetect explains why).
88 enabled units, 89 loaded modules and two authorized_keys files (the lab VM's lima user's key and an empty file for root). After eradication, compare the same three lists again.
diff prints nothing when the lists match, so no output means the enabled units, the loaded modules and the keys are all as they were before the incident: the two det-ir units were enabled after the baseline and are gone again. Expect some noise on a real host: modules load on demand when hardware or a program asks for them, so a new name in lsmod is a question (modinfo <name> shows the file it came from), not yet a finding. Next, the udev and module locations.
/etc/udev/rules.d is empty, as it is on a default Ubuntu server. Every file in /etc/modprobe.d and /etc/modules-load.d belongs to a package (kmod, mdadm, and systemd for the modules.conf link to the old /etc/modules file), no install line exists anywhere, and find -cnewer lists nothing whose inode changed after the baseline's stamp. It compares change times (ctime), which touch cannot set back, unlike modification times (linux-det/hunting). A local udev rule, an unowned modprobe.d file or any install line would each deserve a look.
The same pattern covers every location in the diagram: one command that answers "did this change since the host was trusted, and who owns it". B is the baseline directory above; the commands are the Ubuntu forms, with rpm -qf in place of dpkg -S on RHEL and the platform paths named earlier.
# systemd units and timers: the enabled set, then the audit from linux-det/pedetect# (unit states, unit-file hashes, dpkg --verify of packaged units)systemctl list-unit-files --state=enabled --no-legend | awk '{print $1}' | sort | sudo diff $B/enabled-units -# user units and lingering accountsls /var/lib/systemd/lingersudo find /root /home -path '*/.config/systemd/user/*' -type f# cron and at (atq only when the at package is installed)sudo find /etc/crontab /etc/cron.* /var/spool/cron -type f -cnewer $B/stampsudo atq# shell startup filessudo find /etc/profile /etc/profile.d /etc/bash.bashrc /root/.bashrc /root/.profile /home/*/.bashrc /home/*/.profile -cnewer $B/stamp# authorized_keyssudo find /root /home -name 'authorized_keys*' -type f -exec sha256sum {} + | sort -k2 | sudo diff $B/authorized-keys -# sudoers and PAM configuration, then PAM modules no package ownssudo find /etc/sudoers /etc/sudoers.d /etc/pam.d -cnewer $B/stampfor m in /usr/lib/*/security/*.so; do dpkg -S "$m" >/dev/null 2>&1 || echo "no package: $m"; done# /etc/ld.so.preloadls -l /etc/ld.so.preload# udev rulesls -A /etc/udev/rules.dsudo find /etc/udev/rules.d /usr/lib/udev/rules.d -cnewer $B/stamp# kernel modules loaded at boot, and modprobe install lineslsmod | awk 'NR > 1 {print $1}' | sort | sudo diff $B/modules -dpkg -S /etc/modules-load.d/* /etc/modprobe.d/*grep -rhE '^[[:space:]]*install' /etc/modprobe.d /usr/lib/modprobe.d
Each line either prints nothing or names something to explain. None of them proves the host clean against an intruder who had root (they could have edited the baseline too), which is why the decision above is still a rebuild; they tell you whether the rebuilt host, and every other host that shared the indicators, starts from a state you can vouch for.
Turn it into a detection
The question that pays for the incident is which detection would have caught it on the first minute. The persistence watch on /etc/systemd/system from linux-det/auditpipe would have caught the unit being written. A second rule catches the behaviour whichever persistence method comes next: a program executed from a world-writable directory. auditd is not installed on Ubuntu by default (the lab host has it); on RHEL it runs out of the box.
## Programs executed from world-writable directories-a always,exit -F arch=b64 -F dir=/tmp -F perm=x -F key=exec_from_tmp-a always,exit -F arch=b64 -F dir=/var/tmp -F perm=x -F key=exec_from_tmp-a always,exit -F arch=b64 -F dir=/dev/shm -F perm=x -F key=exec_from_tmp
auditctl -l shows each rule with -S execve added: perm=x on a directory rule means execution, and arch=b64 covers 64-bit programs (aarch64 on this lab, x86_64 on most servers). A listed rule is not yet a working one. Replay the behaviour harmlessly, a copy of true in /var/tmp started by systemd as the implant was, and look for the record.
The record holds the incident's signature: exe=/var/tmp/det-ir-canary, ppid=1 (systemd started it), auid=unset (no login session behind it, so a service) and uid=root. --input-logs makes ausearch read the audit log even when its input is not a terminal, as in a script or a cron job. Expect some legitimate hits, installers in particular unpack and run helpers from /tmp, and tune them out by executable after review rather than deleting the rule.
RHEL caught this one before any rule. On the Rocky lab the same unit never ran.
The file created in /var/tmp is labelled user_tmp_t, and the targeted SELinux policy does not let systemd (init_t) execute that type, so every start fails with status=203/EXEC and an AVC denial while Restart=always retries every two seconds. An intruder with root can relabel the file or move it into a system directory, so treat this as a loud signal (a crash-looping unit and repeated denied { execute } records) rather than as the control you rely on.
Where to go from here
This course followed an intruder across a Linux host and built the defender's answer at each step. You can now map a server's attack surface and the traces each stage of an intrusion leaves, find escalation paths before an attacker does and read containers from the host, feed the audit pipeline, osquery, eBPF and Falco into detections that survive a renamed binary, hunt from a hypothesis, triage a live host without destroying evidence, reconstruct a timeline from its artefacts, and run an incident round this loop.
Two directions follow from here. linux-perf (Advanced Linux internals and tooling) is optional depth on how processes, memory, storage and networking work underneath and how to measure them, including namespaces and cgroups (linux-perf/k-ns) and the eBPF tools (linux-perf/bpftools); it requires only linux-ess. For containers, docker-int (Advanced container security) builds containers from namespaces and cgroups and walks the escape paths and the sandboxes that close them (it assumes docker-fund and docker-hard), and runtime-sec (Runtime and eBPF security) applies eBPF, Falco, Tetragon and Cilium to Kubernetes workloads (it assumes k8s-fund).
Try this
On a lab host with auditd and the rule above loaded, plant your own harmless incident: sudo install -D -m 755 /usr/bin/gnusleep /var/tmp/.probe/probe, then a unit /etc/systemd/system/probe.service with ExecStart=/var/tmp/.probe/probe 600 and Restart=always, and sudo systemctl daemon-reload && sudo systemctl enable --now probe.service. Scope it with systemctl status, systemctl show -p Restart -p NRestarts probe.service and sha256sum /var/tmp/.probe/probe /usr/bin/gnusleep. Capture it before anything else: pid=$(systemctl show -p MainPID --value probe.service), then sudo cp /proc/$pid/exe /root/probe-exe and sudo sha256sum /root/probe-exe. Then kill that PID and confirm NRestarts=1. Contain and eradicate it: stop and disable it, delete the unit and the binary, run daemon-reload, and confirm systemctl status probe.service reports that the unit could not be found. Finally, sudo ausearch --input-logs -k exec_from_tmp -x /var/tmp/.probe/probe -i shows the executions with ppid=1 and auid=unset, and sudo rm /root/probe-exe removes your evidence copy.
Takeaway
Scope before you touch: find what runs it, what revives it and where else it lives. Capture the running process and its files before any kill, stop the reviver with the process, delete every piece, check every persistence location against the package manager or a baseline, and do not close the incident until a detection for it has fired in front of you.
systemctl stop, and 15 seconds later systemctl is-active reports it active again. The unit has Restart=always. What most likely started it?systemctl stop is not restarted, whatever Restart= says.Restart= explains a comeback after a kill, not after a stop, so another unit or job is starting it, and only a search for its name finds that.enabled only means the unit starts at boot through its WantedBy= link; it does not restart a stopped unit./usr/bin/ncat, and dpkg -S finds no package owning its path under /var/tmp. What does that tell you?dpkg -S just showed.dpkg -S searches package file lists by path; a path no package installed is simply not found, whatever the file contains./var/tmp. How do you establish that it would have caught this incident?exe in /var/tmp, ppid=1 and auid=unset is the incident's signature, produced safely on demand.