Threat hunting on Linux
Hypotheses, queries and closing the loop.
Threat hunting is the work of looking for the intrusion your detections never fired on. No rule set is ever complete, so someone has to go looking for the techniques nobody automated yet, form a specific idea of what an attacker would have done, ask the host whether it happened, and turn every answer into a rule so a machine catches it next time. This lesson runs three hunts on one host: a persistence timer nobody scheduled, a process running from a file that was deleted while it ran, and a listener that belongs to no managed service. For each, the lab planted one harmless example (the first section lists exactly what), and you find it with a query keyed on the behaviour, separate it from the benign version of the same signal, and close the loop into a check. It builds on the baselines from linux-det/pedetect and the audit skills from linux-det/auditpipe; the hunts here are the proactive counterpart to those scheduled diffs.
Start with a hypothesis, not a wander
"Look around the servers" has no finish line, so it never proves anything. A hunt starts with one concrete attacker behaviour and the exact trace it would leave, phrased so the host can answer yes or no. "If someone planted persistence with a systemd timer, there is a .timer unit on disk that was written after the host's last known-good baseline, pointing at an odd command." That you can test. Systemd timers are a real persistence technique, catalogued as T1053.006 (Scheduled Task/Job: Systemd Timers) in MITRE ATT&CK, and system units live in /etc/systemd/system and /usr/lib/systemd/system (plus drop-in *.d/*.conf files beside them), while user units live in /etc/systemd/user, /usr/lib/systemd/user and each account's ~/.config/systemd/user. The hunt covers all of them.
The lab planted three examples as root, right after noting the time as the stand-in for the last reviewed baseline. The first is a timer and service named det-hunt-refresh: the service runs /bin/true (its unit is shown below), the timer has OnBootSec=2min and OnUnitActiveSec=10min and was enabled with systemctl enable --now, and the lab then set both unit files' modification time back to 2025-06-12, which anyone with root can do. The other two are a deleted-binary process and a listener, started as deploy with these commands:
install -d -m1777 "/var/tmp/det-hunt"runuser -l deploy -c "cp /usr/bin/gnusleep /var/tmp/det-hunt/.systemd-worker && setsid -f /var/tmp/det-hunt/.systemd-worker 600 </dev/null >/dev/null 2>&1"rm -f "/var/tmp/det-hunt/.systemd-worker"runuser -l deploy -c "setsid -f python3 -m http.server 8081 --bind 127.0.0.1 </dev/null >/dev/null 2>&1"
Hunt one: a timer nobody scheduled
Systemd lists every timer and the unit it starts. That list is an inventory, not a verdict: a planted timer sits in it looking as ordinary as the shipped ones.
Nothing in that line says "malicious". A cheap first filter is age: ask the filesystem which unit files were modified after the last time you reviewed the host.
-newermt compares each file's modification time (mtime) against a reference, here the baseline time, and prints only what is newer (-newer <file> does the same against a reference file, such as one your baseline job writes). Only one entry appears: the symlink in timers.target.wants that systemctl enable created. The service and timer files are missing because their mtime now says 2025, and mtime is exactly the timestamp the owner of a file, or root, can set to any value (utimensat(2); linux-det/artifacts shows what that looks like in stat). Packages muddy it from the other side: dpkg and rpm give installed files the date recorded in the package, so a freshly updated unit can carry an old date. Keep -newermt as a quick first pass and follow it with two checks that do not trust mtime.
The change time (ctime) moves whenever the inode changes, including when someone rewrites the other timestamps, and no ordinary system call sets it back, so all three entries reappear. The query also covers the user unit directories and drop-ins: a foo.service.d/override.conf that replaces ExecStart= turns a trusted unit into persistence without creating any new .service file. ctime narrows the field rather than deciding it, because a package upgrade, a chmod or a chown moves it too. What decides is provenance: a unit file that no package installed arrived by some other route.
The loop asks dpkg which package owns each unit file. Some Ubuntu packages still register their units under /lib/systemd/system, the path from before the /usr merge (/lib is now a link to /usr/lib), and dpkg -S matches the path as the package recorded it, so the loop retries that spelling before reporting a file. Four files are left. Two are the plant. lima-guestagent.service and the user@.service.d/lima.conf drop-in belong to Lima, the tool that runs this lab VM, which is the point: on a real host every unowned unit is either something your team wrote and can account for, or a finding. dpkg -S says who owns a path, not whether its content still matches the package; dpkg --verify (Ubuntu) and rpm -Va (RHEL) answer that. On Rocky Linux the packages record /usr/lib paths, so rpm -qf needs no second try.
Root is not required for this kind of persistence. Any account can install user units under ~/.config/systemd/user, and with lingering enabled (loginctl enable-linger, which leaves a file named after the account in /var/lib/systemd/linger) its user manager starts at boot and runs them without a login. A foothold in a service account can do exactly that, so the hunt reads the linger list and the home directories too.
The linger list holds only lima, the VM tool's own account, and no home directory has user units. On your hosts, compare this list with the accounts that are supposed to run user services. The unit database is the next place to look, though what it tells you depends on the distribution.
STATE says enabled and PRESET, the column that records what the vendor's preset policy would do with the unit, also says enabled. On Ubuntu that is no help: Ubuntu ships no catch-all preset rule, and systemd treats a unit that no preset line matches as enable (systemd.preset(5)), so every planted unit looks vendor-approved here. RHEL ships the opposite default, a final disable * line, and the same planted unit tells on itself there.
On Rocky Linux the unit is enabled against a disabled preset, which means someone turned it on by hand rather than the distribution. That is not proof of compromise, because ordinary local changes look the same, but on RHEL it narrows the field to units a person chose to enable. On Ubuntu you rely on provenance and change time instead. Either way, read what the service actually runs.
In this lab the ExecStart is /bin/true, a harmless stand-in. On a real host this is where you find the beacon: a curl to a bare IP piped into a shell, a script in a hidden path, a reverse connection. A timer with OnBootSec and OnUnitActiveSec that runs something unexplained on a schedule is persistence wired to survive reboots. The same find query, pointed at every host, becomes the fleet hunt.
Hunt two: a process whose file is gone
Attackers like to run from memory and leave nothing on disk to scan. A payload copies itself into a scratch directory, starts, then deletes its own file; the program keeps executing while the file that launched it no longer exists. Deleting a running program's file to cover tracks is technique T1070.004 (Indicator Removal: File Deletion). The kernel keeps a link to the original binary at /proc/<pid>/exe even after the file is unlinked, and it marks it, so one listing surfaces every process running from a deleted or scratch-directory file.
One process is running from /var/tmp/det-hunt/.systemd-worker, and the kernel has appended (deleted) because the file was removed after the process started. A hidden name (the leading dot), a scratch directory, and a backing file that no longer exists: that combination is a strong, low-noise lead. Pivot to who and what.
The process runs as deploy from a deleted file, and its parent PID is 1. Services that systemd starts also show PPID 1, but no service owns this process: whatever started it has exited, so it was adopted by init and the lineage in ps is gone (the audit log from linux-det/auditpipe keeps the original parent). Because the kernel still maps the bytes, you could copy the binary straight out of /proc/<pid>/exe for analysis before the process exits; linux-det/triage does exactly that during live response. Cryptominers such as the widely reported kdevtmpfsi disguise themselves as kernel threads and behave this way, so a real hit here is often the start of an incident.
(deleted) for a dull reason: its package was upgraded and the old binary was replaced while the process kept running the version already in memory. sshd the morning after a patch window is the classic false positive (linux-ess/processes noted this). That is why you triage every hit by path, parent and owner before paging anyone. A daemon running a deleted binary from /usr/sbin after an upgrade is routine; a hidden file in /var/tmp with a deleted backing store is not. The flag is a lead, never a verdict.Hunt three: a listener that maps to no service
Anything an attacker uses to receive connections has to open a socket, and every listening socket has a process behind it. List them with the owning program attached.
A default server is quiet: sshd on port 22 and systemd-resolved's stub resolver on the loopback addresses 127.0.0.53 and 127.0.0.54 (linux-det/threatmodel walked this set). The line that does not belong is the python3 process listening on 127.0.0.1:8081, which no packaged service placed there. A listener you cannot trace back to a managed unit is a question, whether it binds the loopback address (reachable only from the host, often a local relay or a foothold's staging port) or a public one. The benign version of both of the last two signals is worth naming out loud.
python3 itself is packaged and legitimate: the interpreter is not the problem, the unmanaged listener it was told to run is. A package install creates new unit files and a package upgrade leaves deleted-binary processes, so all three hunts turn up benign hits constantly; the judgement is always path, owner and provenance, not the raw signal.
Hand off before you clean up
Three findings on one host (a hand-made timer with backdated files, a hidden process running from a deleted file, an unmanaged listener) add up to a probable intrusion, and at that point the hunt stops being a hunt. Removing anything now would destroy evidence: kill the process and its /proc/<pid>/exe link, the only copy of the deleted binary, goes with it. Capture first, into a root-only evidence directory: the executable out of /proc, the process details, the listener table, the unit files with their timestamps, and a hash of each (linux-det/triage covers the full live capture).
The fifth line is the hash of the binary recovered from /proc; unit-times.txt records the backdated mtime next to a ctime from today, which is itself evidence of tampering. Next comes scoping: search the other hosts for the same hash, path and unit names before touching this one, because cleaning one host first tells the intruder what you can see (linux-det/ir runs that loop). Containment and eradication follow from there. In the lab, the teardown below stands in for that later eradication, and a check confirms the host is quiet again.
A hunt that pays off ends by leaving something behind: the queries become a scheduled check that records the shape of a healthy host today (linux-det/pedetect builds exactly that baseline-and-diff), so the same technique is caught automatically next time and you spend the afternoon hunting what nobody has automated yet. A hunt that finds nothing is fine, but only if it ends in a new detection or a sharper baseline.
Try this
Run the timer hunt end to end on a lab host. Note the current time as your baseline, then create a unit det-x.service (Type=oneshot, ExecStart=/bin/true) and det-x.timer (OnUnitActiveSec=10min) in /etc/systemd/system, and systemctl enable --now det-x.timer. Now hunt: run sudo find /etc/systemd/system /usr/lib/systemd/system \( -name '*.timer' -o -name '*.service' \) -newerct '<your baseline time>' and confirm your two files and the enable link appear, then run dpkg -S /etc/systemd/system/det-x.timer (rpm -qf on RHEL) and confirm no package owns it. Run systemctl list-unit-files det-x.timer and note the PRESET column (enabled on Ubuntu, disabled on RHEL and Rocky). Remove both units, systemctl daemon-reload, and re-run the find to confirm the host is clean. You have just done, by hand, the hunt that a scheduled check would run for you.
Takeaway
Hunt from a written hypothesis about one attacker behaviour and the trace it leaves, then judge every hit by path, owner and provenance rather than the raw signal or a timestamp the attacker can set, because the same signals, a new unit, a deleted binary, an unexpected listener, all have common benign causes. Capture before you clean, and end every hunt by turning it into a check.
systemctl list-timers and a suspicious sysupdate.timer sits there looking as ordinary as apt-daily.timer. Which comparison actually separates a planted unit from one the distribution shipped?touch; together they survive a backdated plant./proc hunt turns up two processes whose /proc/<pid>/exe points at a deleted file: sshd from /usr/sbin the morning after a patch window, and a hidden file in /var/tmp owned by a service account, holding an outbound connection. How should you handle them?/var/tmp process keep beaconing; the point of triage is to keep the real hit while dismissing the routine one.sshd row, while nothing routine explains a hidden deleted binary in /var/tmp with a live outbound socket./usr/sbin is routine, while a hidden deleted binary in a scratch directory beaconing out fails all three checks.