Parse auditd logs with Python before they hit your SIEM
Raw audit records are unreadable and expensive to index. Normalize them with 80 lines of Python and cut ingest cost.
grep "audit(1757667721.324:8841)" /var/log/audit/audit.logtype=SYSCALL msg=audit(1757667721.324:8841): arch=c000003e syscall=257 success=yes exit=3 … auid=1000 uid=0 gid=0 euid=0 … comm="vim" exe="/usr/bin/vim" key="identity"type=CWD msg=audit(1757667721.324:8841): cwd="/root"type=PATH msg=audit(1757667721.324:8841): item=0 name="/etc/" inode=2 … nametype=PARENTtype=PATH msg=audit(1757667721.324:8841): item=1 name="/etc/passwd" inode=8391 … nametype=NORMALthe question "who edited /etc/passwd" needs fields from three of these four linesThe audit log is precise and hostile. One logical event, a process opening a file, is written as several records that share a serial number and are otherwise independent lines of key=value pairs: the syscall record has the identity (auid is the login user, uid is who they became), the CWD record has the directory, and one PATH record per path component has the file. ausearch -i joins them for a single host and turns numbers into names. When the question spans a fleet, or the answer has to land in Loki or a SIEM as one object per event, forty lines of Python do the same join and produce JSON.
The shape of the records
| Record | Carries | Field to keep |
|---|---|---|
SYSCALL | who, what call, whether it succeeded, which binary | auid, uid, syscall, success, exe, comm, key |
CWD | the working directory the path is relative to | cwd |
PATH (one per item) | each path the syscall touched: parent directory, the file, a rename target | name, nametype (NORMAL, CREATE, DELETE, PARENT) |
PROCTITLE | the command line, hex-encoded when it contains spaces | proctitle (decode from hex) |
EXECVE | argv of an executed program, one field per argument | a0, a1, … |
Join by serial, keep every PATH
Two details decide whether the parser is right. Values are sometimes quoted (name="/etc/passwd") and sometimes not (auid=1000), so the field regex has to accept both. And a single event carries several PATH records, so the parser must collect them as a list instead of letting the last one overwrite the rest; the record that matters is usually the one with nametype=NORMAL or CREATE, not the PARENT directory that comes first.
import collections, json, re, sysFIELD = re.compile(r'(\w+)=(?:"([^"]*)"|(\S+))')HEAD = re.compile(r'^type=(\w+) msg=audit\((\d+\.\d+):(\d+)\):')UNSET = "4294967295"def parse(lines):"""Yield one dict per audit event, joining every record that shares a serial."""events = collections.OrderedDict()for line in lines:head = HEAD.match(line)if not head:continuertype, ts, serial = head.groups()fields = {k: q or u for k, q, u in FIELD.findall(line[head.end():])}ev = events.setdefault(serial, {"serial": serial, "time": float(ts), "paths": []})if rtype == "PATH":ev["paths"].append({"name": fields.get("name"), "nametype": fields.get("nametype")})elif rtype == "SYSCALL":ev.update({k: fields.get(k) for k in ("auid", "uid", "syscall", "success", "exe", "comm", "key")})elif rtype == "CWD":ev["cwd"] = fields.get("cwd")return events.values()if __name__ == "__main__":for ev in parse(open(sys.argv[1] if len(sys.argv) > 1 else "/var/log/audit/audit.log")):print(json.dumps(ev))
The OrderedDict keyed by serial is the whole join; records for one event are adjacent in practice but the parser does not rely on it. Everything else is choosing which fields to keep, and the list is deliberately short: a SIEM ingest bill is proportional to what you ship, and the fields above answer the identity questions without the fifty others a SYSCALL record carries.
Ask the question
from parse_audit import parse, UNSETfor ev in parse(open("/var/log/audit/audit.log")):if ev.get("key") != "identity" or ev.get("auid") in (None, UNSET):continue # not a watched file, or not a logged-in humantarget = next((p["name"] for p in ev["paths"] if p["nametype"] in ("NORMAL", "CREATE", "DELETE")), None)print(f'{ev["time"]:.0f} auid={ev["auid"]} as uid={ev["uid"]} {ev["comm"]} -> {target} ({ev["success"]})')
python3 who.py1757667721 auid=1000 as uid=0 vim -> /etc/passwd (yes)1757671104 auid=1000 as uid=0 useradd -> /etc/passwd (yes)login user 1000 became root twice and edited the file; auid survives sudo, uid does not. The daemon write (auid unset) and the network syscall (a different key) are in the log but not in this listausearch -k identity -ts today -i | grep -E "^type=(SYSCALL|PATH)" | grep -E "auid=|nametype=NORMAL"type=SYSCALL … auid=alice uid=root … comm=vim exe=/usr/bin/vim key=identitytype=PATH … name=/etc/passwd … nametype=NORMALfor one host and one question, ausearch -i already joins and interprets; the script earns its keep at fleet scaleauid is the field that makes the answer meaningful: it is set at login and inherited through su and sudo, so a change made as root still names the person. The sentinel 4294967295 (unset) marks processes with no login session, which is every daemon, and filtering it out is what turns the output into a list of humans. ausearch --format csv or --format text is the no-code route to a readable export, and -i is the flag that turns auid=1000 into auid=alice on the host that knows the mapping.
Ship one object per event
Printing JSON lines is the bridge to everything downstream: an Alloy or Vector pipeline tails the output (or the parser runs as a sidecar), the key becomes a label, auid becomes structured metadata, and "who touched /etc/shadow on any host this month" is a saved query rather than an ssh loop. For a live stream instead of a file, journalctl -u auditd -o cat or the audit dispatcher's syslog plugin provide the same lines on stdin, and parse() does not care which.
The rules that generate these events, with a key per question and the filters that keep the volume readable, are in auditd rules that satisfy the auditor. Where the JSON lands and how it is queried next to application logs is in logs in Loki.
Go deeper in a coursePython for security automationLog parsing, APIs and small CLI tools with real error handling.View course