Centralized logging with journald and rsyslog
Make the journal persistent, forward it to a central host, and keep logs you can actually search after an incident.
journalctl -b -1 -u ssh --since "-2h"Specifying boot ID or boot offset has no effect, no persistent journal was found.ls -d /var/log/journal 2>/dev/null || echo "no persistent journal directory"no persistent journal directorythe host rebooted during the incident and everything before the reboot is goneOn a systemd host, journald receives everything: kernel messages, service stdout, PAM and sshd events, and the audit dispatcher if you route it there. Whether any of that survives a reboot depends on one setting and one directory, and on many cloud images the default is neither. A journal that lives in /run is a debugging convenience; a journal on disk, capped, rate-limited and forwarded, is evidence.
Persist it, cap it, and keep it from drowning
[Journal]Storage=persistent # write to /var/log/journal (created on restart if missing)Compress=yesSystemMaxUse=2G # cap on disk; oldest files are rotated out firstSystemKeepFree=1G # never eat the last gigabyte of the filesystemMaxRetentionSec=1monthRateLimitIntervalSec=30s # per-service burst limit: a looping daemon cannot crowd out sshdRateLimitBurst=10000Seal=yes # forward secure sealing: tampering with sealed files is detectableForwardToSyslog=no # rsyslog reads the journal directly (imjournal); avoid duplicates
SystemMaxUse and SystemKeepFree are the two limits that decide whether a chatty service or a small root volume takes the host down; the smaller of the two wins. The rate limit is per service (per cgroup), so a service that logs a stack trace ten times a second is throttled while authentication events from sshd keep arriving, which is the property you want during the exact incident that produces the flood. Seal=yes with a verification key makes silent editing of the on-disk files detectable through journalctl --verify; it does not stop deletion; forwarding covers that.
systemctl restart systemd-journald && journalctl --disk-usageArchived and active journals take up 48.0M in the file system.ls -ld /var/log/journal/*/ && journalctl --header | grep -E "^(State|Sealed)"drwxr-sr-x 2 root systemd-journal 4096 Sep 12 09:02 /var/log/journal/3f1e…/State: ONLINESealed: YESbake this into the image; a first-boot playbook that fails leaves the default behindThree ways off the host
Forwarding options
| Path | How | Fits when |
|---|---|---|
rsyslog, imjournal in, omfwd out over TLS | rsyslog reads the journal natively and forwards RFC 5424 with the gtls stream driver; the collector runs imtcp with TLS or is a SIEM syslog receiver | the destination speaks syslog; you already run rsyslog |
systemd-journal-upload to systemd-journal-remote | HTTPS push of journal entries with all fields intact; the collector stores real journal files you can query with journalctl -D | systemd on both ends, and you want the structured fields, not a flattened line |
| an agent (Grafana Alloy, Vector, Fluent Bit) reading the journal | the agent tails the journal, keeps the cursor, adds labels, ships to Loki, OpenSearch or object storage | the destination is a log platform with its own query language |
module(load="imjournal" StateFile="imjournal.state") # read the journal directlyglobal(DefaultNetstreamDriverCAFile="/etc/rsyslog/ca.pem")action(type="omfwd"target="logs.acme.dev" port="6514" protocol="tcp"StreamDriver.Name="gtls" StreamDriver.Mode="1"StreamDriver.AuthMode="x509/name" StreamDriverPermittedPeers="logs.acme.dev"queue.type="LinkedList" queue.filename="fwd" queue.saveOnShutdown="on"action.resumeRetryCount="-1")# plain @@host:6514 would send cleartext to the TLS port; the StreamDriver lines are the TLS
Whichever path, two properties matter more than the transport: a disk-backed queue on the sending side so a collector outage does not drop events (queue.filename above; Alloy and Vector have equivalents), and local retention that outlives the outage, because the collector fails during the same incidents you need it for. Forwarding replaces nothing on the host; it adds a copy an attacker with root on the host cannot delete.
The queries that answer incident questions
journalctl -u ssh -S "2026-09-11 22:00" -U "2026-09-12 02:00" | grep -E "Accepted|Failed"Sep 11 23:41:07 app-14 sshd[8821]: Accepted publickey for deploy from 10.0.4.22 port 51842 ssh2: ED25519-CERT …journalctl -b -1 -p err..alert --no-pager | tail -5errors from the previous boot: this is the query that needs Storage=persistentjournalctl _UID=0 _COMM=bash --since today -o json | jq -r ".MESSAGE" | headfield matches (_UID, _COMM, _EXE, _SYSTEMD_UNIT) are indexed; grep on the text is notjournalctl -u api --since "-15m" -o json-pretty | head -40 > /tmp/api-context.jsonjson-pretty keeps the metadata a ticket attachment needs and a copy from the terminal losesTwo numbers are worth an alert once this is in place. The first is the journal's own disk usage against SystemMaxUse: when the cap is reached, rotation starts deleting the oldest files, which during a long incident means deleting the beginning of it, so a gauge that shows the journal near its cap is a prompt to raise the cap or lower retention before the evidence goes. The second is the forwarder's queue depth (the rsyslog queue.filename file growing, or the agent's equivalent metric): a queue that keeps growing means the collector has been unreachable for longer than anyone noticed, and the host is the only place the last hours exist.
The journal is one of two host streams worth centralising; the other is the audit trail from auditd rules, which the audit dispatcher can also hand to journald. Where both end up, and how they are queried together, is the subject of logs in Loki.