Finding escalation paths before an attacker does
Baseline and diff the risky surfaces.
The last three lessons showed how local root is reached: SUID files, sudo rules, capabilities, unpatched kernels, weak jobs, leaked secrets. This lesson turns that knowledge into a job you run on your own hosts. The method is the same for every surface: enumerate it, compare it with a reviewed baseline, and treat every addition as a finding. You will build a small audit script that snapshots eight surfaces, put it behind a systemd timer that fails loudly when something changes, watch it catch a planted SUID file, a new sudoers rule, a new timer and an edited packaged unit in one run, and close each finding without destroying its evidence. You will also see what an on-host detector cannot protect against, and what covers that gap.
Enumerate every real filesystem, and know your noise
A scan that drowns you in noise is a scan you will stop reading. World-writable files are the clearest example. Searched naively, the count is meaningless.
Every hit is under /proc and /sys, the kernel's virtual filesystems, where a world-writable mode is a tunable knob, not a file on disk. (The two counts differ because /proc changes between runs; each process adds its own entries.) Two changes fix it. -xdev keeps find on the filesystem it started on, so it never descends into /proc, /sys, /run or a network mount. For directories, add the sticky-bit test: a world-writable directory with the sticky bit set (/tmp, /var/tmp) is intentional, because the sticky bit stops users deleting each other's files.
Zero, the right answer for a clean root filesystem. The price of -xdev is that it stops at mount points, so a single scan of / misses every other on-disk filesystem, and an intruder with root can drop a SUID file on any of them. Ask findmnt for the real ones instead of guessing.
This host has two, / and /boot; RHEL's automatic partitioning adds a /home file system when the disk has at least 55 GiB, and hardened layouts commonly split out /var as well. /tmp is a tmpfs mounted nosuid,nodev, so a SUID file there is inert (linux-det/peloc showed this) and it is left out. The script below loops over exactly this list, so a new volume is covered the day it is mounted.
One script, eight surfaces, a reviewed baseline
The surfaces worth watching are the ones the earlier lessons abused, plus the places persistence usually lives. SUID and SGID files are recorded with a SHA-256 hash, so a trojaned passwd changes a line even though its name did not. Capabilities are read file by file with find ... -exec getcap, because getcap -r / has no -xdev and would walk /proc, /sys and any network mount; each line also gets the file's hash, so a replaced ping that kept its cap_net_raw still shows up. sudoers rules include /etc/sudoers-rs: when that file exists, sudo-rs uses it in place of /etc/sudoers (the sudo-rs README), so an attacker could replace the whole policy without touching /etc/sudoers; also check that no @includedir points somewhere else. Units record the enabled state of every service and timer, and unit files hash everything under /etc/systemd/system, drop-ins included, so an edited ExecStart= shows up. Packaged units in /usr/lib/systemd/system change with every update, so hashing them would alert on each upgrade; the packaged-units surface asks the package manager instead. dpkg --verify compares every packaged file with the checksum its package recorded and prints only the files that differ (rpm -Va does the same on RHEL), and the script keeps the lines under /usr/lib/systemd (and /lib/systemd, the path a few Ubuntu packages still record). An upgrade installs new checksums with the new files, so only an in-place edit appears. Cron covers /etc/crontab, /etc/cron.d, /etc/anacrontab and the whole of /var/spool/cron (Debian keeps user crontabs in /var/spool/cron/crontabs, cronie on RHEL in /var/spool/cron/<user>), plus hashes of the scripts in the cron.hourly to cron.monthly directories. The last surface is world-writable files and directories.
#!/bin/sh# pd-privesc-audit: snapshot the local privilege-escalation surface and# compare it with a reviewed baseline. Prints each added (+) or removed (-)# line and exits 1 when anything changed.# Accept the current state as the new baseline: pd-privesc-audit --acceptset -udir=/var/lib/pd-privescumask 077mkdir -p "$dir/now" "$dir/baseline"cd "$dir/now" || exit 2# Every on-disk filesystem; -xdev keeps find off /proc, /sys, /run and network mounts.mounts=$(findmnt -rno TARGET -t ext4,xfs,btrfs)find $mounts -xdev \( -perm -4000 -o -perm -2000 \) -type f -exec sha256sum {} + 2>/dev/null | sort -k2 > suid-sgidfind $mounts -xdev -type f -exec getcap {} + 2>/dev/null |while read -r f caps; do echo "$(sha256sum < "$f" | cut -c1-64) $f $caps"; done | sort -k2 > capabilitiesgrep -rvE '^[[:space:]]*(#|$)' /etc/sudoers /etc/sudoers-rs /etc/sudoers.d 2>/dev/null | sort > sudoerssystemctl list-unit-files --type=service,timer --no-legend | sort > unitsfind /etc/systemd/system -type f -exec sha256sum {} + 2>/dev/null | sort -k2 > unit-files# Packaged systemd files changed in place: dpkg --verify (Debian, Ubuntu) or rpm -Va (RHEL)# compares every packaged file with the package's checksum; keep the systemd lines.if command -v dpkg >/dev/null; then dpkg --verify; else rpm -Va; fi 2>/dev/null |grep -E ' (/usr)?/lib/systemd/' | sort > packaged-units{ grep -rvE '^[[:space:]]*(#|$)' /etc/crontab /etc/cron.d /etc/anacrontab /var/spool/cronfind /etc/cron.hourly /etc/cron.daily /etc/cron.weekly /etc/cron.monthly -type f -exec sha256sum {} +} 2>/dev/null | sort > cronfind $mounts -xdev -perm -0002 \( -type f -o -type d ! -perm -1000 \) 2>/dev/null | sort > world-writableif [ "${1:-}" = --accept ]; thencp ./* ../baseline/ && echo "baseline accepted: $(ls | wc -l) surfaces" && exit 0fi[ -e ../baseline/suid-sgid ] || { echo "no baseline yet: review $dir/now, then run with --accept"; exit 2; }changed=0for f in *; doif ! cmp -s "../baseline/$f" "$f"; thendiff "../baseline/$f" "$f" | sed -n "s/^> /$f: + /p; s/^< /$f: - /p"changed=1fidone[ "$changed" = 0 ] && echo "no change since the baseline"exit "$changed"
The first run has nothing to compare against, so it saves the current state and asks you to review it before you promise it is clean. It runs from a oneshot service at low CPU and I/O priority (Nice=, IOSchedulingClass=idle), started by an hourly timer.
[Unit]Description=Compare the privilege-escalation surface with its baseline[Service]Type=oneshotExecStart=/usr/local/sbin/pd-privesc-auditNice=10IOSchedulingClass=idle
[Unit]Description=Hourly privilege-escalation surface check[Timer]OnCalendar=hourlyRandomizedDelaySec=5minPersistent=true[Install]WantedBy=timers.target
Save the three files as root, make the script executable and reload systemd, then run the script once and read what it captured.
Eight files, one per surface: 22 SUID/SGID binaries, four capability-bearing files, ten effective sudoers lines, 308 service and timer unit files with their state, 22 cron lines and script hashes, and no packaged systemd file and no world-writable file that is out of place. The four unit files are the audit's own two plus two that the lab VM's tooling (Lima) installs; a plain Ubuntu server has neither of those. Reading this list once, by hand, is the whole point of a baseline: you certify that every line belongs before you start alerting on change. Enable the schedule, then accept the reviewed state so the timer you just enabled is part of it.
RandomizedDelaySec spreads a fleet's runs so they do not all start on the hour, and Persistent=true catches up a run missed while the host was off (linux-ess/cron). Accepting after enabling keeps the audit's own timer from showing up as a change on every later run.
Who the baseline protects against
umask 077 and the root-only /var/lib/pd-privesc stop unprivileged accounts from reading the baseline (to learn what you trust) or editing it (to hide a change). They do nothing against root, and every change this detector exists to catch, a new SUID-root file, a sudoers rule, a system timer, needs root. Root can rewrite the baseline, run --accept, edit the script, or mask the timer, and can rewrite the checksums dpkg --verify compares against, which live on the same host under /var/lib/dpkg/info. So the on-host copy defends against drift and against non-root tampering, and three things cover the rest. Keep the reviewed baseline's hashes off the host, so a quietly rewritten baseline is visible.
Store those eight lines where the host cannot write, such as your configuration repository or the change ticket that approved the baseline. Forward the journal off the host, so each run's findings leave before an intruder can erase them. And alert centrally when a host's hourly result stops arriving, not only when a run fails: a detector that has been switched off looks exactly like a clean host unless something expects to hear from it.
Make the change fail loudly
Now make four changes of the kind the detector must catch, standing in for an intruder with root or an unreviewed change: a SUID-root copy of true, a narrow-looking sudoers drop-in for a new system account (checked with visudo -cf and installed at mode 0440, as a real one would be), and a new timer.
The timer is two small unit files that run /usr/bin/true once a day.
[Unit]Description=Lab stand-in for an unexplained timer[Service]Type=oneshotExecStart=/usr/bin/true
[Unit]Description=Lab stand-in for an unexplained timer[Timer]OnCalendar=daily[Install]WantedBy=timers.target
The fourth change edits a unit that a package installed: fstrim.service from util-linux, whose ExecStart= now runs /usr/bin/true instead of fstrim. Its name, its enabled state and /etc/systemd/system are all unchanged, which is exactly what an intruder editing a stock unit is counting on.
Because the check exits non-zero when a surface changed, the oneshot service that runs it fails, and a failed unit is something monitoring already watches. Start it the way the timer would.
The journal holds the run with the exact lines that differed, each prefixed with its surface: dpkg --verify's line for the edited fstrim.service (5 in the third column means the checksum no longer matches the package), the new sudoers rule granting pd-temp a command as root, the SUID file with its hash, both new unit files with their hashes, and the new timer and its service in the unit list. None of these is loud on its own; the diff surfaces them together. The last line is the price: about 15 seconds of CPU per run, nearly all of it dpkg --verify reading every packaged file, and a memory peak that counts the page cache those reads filled. That is why the service runs at idle I/O priority once an hour, not every minute. The failed unit also shows where operators already look.
--failed lists every failing unit on the host; here the audit is the only one, but on a busy host triage the whole list rather than assuming one cause.
Capture each finding, then close it
A SUID-root file, a sudoers rule, a timer and an edited system unit that no change record explains are an incident, not drift, because only root could have made them. Capture before you remove anything: removing a file loses its content, and changing it updates the change time you would want later.
All three new files are owned by root and were born within a second of each other, which dates the change. The SUID file's hash equals /usr/bin/true's, so it is a renamed copy of a packaged binary, and dpkg -S says no package owns any of the three paths. The edited unit is the opposite case: util-linux owns it, and dpkg --verify util-linux confirms its contents no longer match that package (the runuser-l line above it is the lab VM's own provisioning change, explained in linux-det/ir). Its hash identifies the edited copy; save the file itself before you restore it. On a real host you would copy them to your evidence store, look up who made the change in the audit log (linux-det/auditpipe watches these paths) and follow the incident process (linux-det/triage, linux-det/ir). Closing a finding means removing it, as here, or accepting it into the baseline after review. A packaged file is restored by reinstalling its package, which writes the packaged copy back. Re-running --accept is a reviewed decision, never the way to make an alert go away.
apt-get install --reinstall util-linux fetched the same version again and unpacked it over the installed one, and dpkg --verify util-linux now reports only the lab's runuser-l line. (On RHEL, dnf reinstall util-linux does the same.) On Ubuntu a reinstall keeps configuration files you changed, as the runuser-l line shows, so it does not restore those; for them, compare with the baseline and the package's default by hand.
All four changes are gone and the service finishes successfully, so the alert clears. Tools automate parts of this loop: osquery (linux-det/osquery) exposes these surfaces as SQL tables and reports only what changed, and Lynis grades hardening and flags escalation paths, though its scores vary by host and version, so track findings by test id rather than the number. Audit rules watch the same paths as they change (linux-det/auditpipe); the hourly diff catches slow drift and anything the rules missed, the watch catches the moment of change.
Try this
Prove the detector fires before you trust it. With the baseline in place, plant a harmless SUID decoy: sudo install -m 4755 -o root /usr/bin/true /var/tmp/.probe. Run sudo pd-privesc-audit and confirm it prints a suid-sgid: + line with a hash and /var/tmp/.probe, and exits non-zero. Remove it with sudo rm /var/tmp/.probe, run the check again, and confirm it reports no change. If the first run comes back clean, check that the decoy landed on a filesystem findmnt -t ext4,xfs,btrfs lists (a tmpfs such as /tmp is not scanned) and that the baseline was accepted. To remove the detector afterwards, disable the timer and delete the script, the two unit files and /var/lib/pd-privesc.
Takeaway
Baseline the escalation and persistence surfaces with hashes on every real filesystem, alert on the diff and on a missing result, and keep the baseline's hashes off the host. A privileged file, rule or unit that was not there yesterday is a finding to capture before you clean it.
find / -perm -0002 -type f returns over two thousand hits, all under /proc and /sys. What is the right fix before you baseline it?-xdev stops find entering them in the first place.chmod on them is meaningless or harmful, and they are not an escalation surface.-xdev keeps find out of /proc and /sys, the sticky-bit test drops intentional directories like /tmp, and looping over findmnt's list covers /boot, /home or /var volumes.Permission denied messages, not the /proc and /sys entries, which are genuinely world-writable pseudo-files.pd-privesc-audit.timer as a change. Why, and what is the fix?Persistent= controls catch-up of missed runs; it has nothing to do with what the enumeration records or what the diff reports./var/lib/pd-privesc. What does that result tell you?--accept, change the script or mask the timer; the directory's mode only keeps unprivileged accounts out.