Logging and log drivers
json-file, local and journald, rotation, remote drivers and the dual-logging cache.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-hard/logging.tar.gz && tar -xzf logging.tar.gz, which creates ~/lab/logging/. SHA-256: 8e828b5b791bba4f9060faf6da2f8b22cb69e9144d12d6599c39f9197e3303aaubuntu is not in the docker group and uses sudo. If that VM does not exist yet, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, and open a shell in it with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec with Lima). ./setup/create-lab.sh --profile sec --recreate resets it.A node alerts at 95% disk. The largest file on it is not an image or a volume but /var/lib/docker/containers/<id>/<id>-json.log, the log of a service that has been printing a retry message in a loop since Friday. Nothing rotated it because, by default, nothing does. Every line a container writes to stdout or stderr goes through a logging driver chosen when the container was created, and that driver decides where the bytes go, how much is kept, and whether docker logs can read them back. This lesson compares the drivers you will meet on Linux hosts and shows how each one behaves under load and when its destination is down.
json-file, the default
With no daemon.json, the default driver is json-file and the container's LogConfig is empty: no max-size, no max-file, so the file grows for as long as the container exists. Each line becomes one JSON object with the text, the stream it came from and a nanosecond timestamp, stored under the daemon's data root (still /var/lib/docker/containers/, also with the containerd image store). Docker copies stdout and stderr in separate goroutines, so lines from the two streams are not guaranteed to land in the file in the order the process wrote them. Here they happen to agree: started (stdout) has the earlier timestamp and comes first. When you correlate the streams, rely on the time field, not on line order. Reading the file directly needs root, which is why docker logs (which "Inspect, logs and exec" in Docker for beginners introduced) is the normal interface.
The same output, three ways of storing it
Three containers print the same 200,000 lines: one on plain json-file, one on json-file with rotation and compression, one on the local driver. docker wait returns when all three have exited, printing each exit code.
About 18 MB for 200,000 short lines, roughly 90 bytes per line, most of it JSON wrapping and the timestamp. Scale that to a chatty service over a few weeks and you have the opening incident. Rotation is two options: max-size caps one file, max-file caps how many files are kept, the live one included. compress=true gzips the rotated ones.
The capped container kept one live file under 1 MB and two gzipped rotations of roughly 60 KB each, so the worst it can ever use is about 1 MB plus the compressed history. docker logs reads across the rotated files, including the compressed ones, and returns the most recent 30,244 lines; the oldest surviving line is id 169757, and everything before it is gone for good. Sizing rotation therefore sets two things at once, the disk a container may use and the history an investigator gets. Your file sizes and counts will differ a little from run to run, because rotation happens when a write crosses the size limit and gzip output varies.
The local driver stores the same 200,000 lines in about 10 MB, a little over half of json-file, using a compact binary format, and docker logs still returns all 200,000. Its empty Config hides defaults that json-file lacks: local rotates at 20 MB per file, keeps 5 files and compresses rotated ones, so it is bounded without any options. Docker's documentation recommends local to avoid json-file's unbounded growth. The reason many hosts stay on json-file anyway is tooling: log shippers such as Fluent Bit or Vector are often configured to tail *-json.log files directly, and they cannot read local's format. If yours does, json-file with max-size and max-file is the right default; if your shipper reads the journal or receives logs over the network, local is the better on-host format.
journald
On a systemd host the journald driver writes each line into the system journal with the container's identity attached as journal fields.
The driver adds fields to every entry: CONTAINER_NAME, CONTAINER_ID and SYSLOG_IDENTIFIER with the short ID, and CONTAINER_ID_FULL with the full one, which is what the first two commands filter on. Filtering on the name is tempting, but the journal outlives containers: the count shows how many entries named lab-journal it holds, and on a VM where you ran this section before, that includes the earlier containers' lines. Filter on the ID when you want one container. Reading the journal worked without sudo because ubuntu is in the adm group on the lab VM; on other distributions the group is systemd-journal. Size limits and retention come from journald.conf (SystemMaxUse and friends), and so does rate limiting: a container that floods the journal can have lines dropped by journald, which you raise in journald.conf, not in Docker. docker logs reads the journal back, so the familiar command keeps working.
Remote drivers and the dual-logging cache
Drivers such as gelf, syslog, fluentd, splunk and awslogs send each line to a collector. To see what a collector receives, start a stand-in: a Python script in the lesson files that prints every UDP datagram arriving on port 12201, the usual GELF port.
# A stand-in for a log collector: prints every UDP datagram it receives on port 12201.# GELF messages arrive as one JSON object per datagram when the sender disables compression.import socketsock = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)sock.bind(("0.0.0.0", 12201))while True:data, _ = sock.recvfrom(65535)print(data.decode(errors="replace"), flush=True)
The collector received a GELF JSON message with the line in short_message and the container name and host as extra fields (compression is turned off here so the datagram is readable). docker logs lab-gelf also printed the line, although the gelf driver itself cannot read anything back. Since Engine 20.10, Docker keeps a local copy for drivers that cannot be read, the dual-logging cache:
container-cached.log is that copy, written with the local driver's format and bounded by default at 5 files of 20 MB per container (cache-max-size, cache-max-file, cache-compress change it). With cache-disabled=true, docker logs fails with configured logging driver does not support reading. So the claim "with a remote driver, docker logs shows nothing" is outdated; it is true only when someone disabled the cache. Disable it deliberately when the extra local copy is a problem, for example because the logs contain data that must not stay on the node.
Transport matters when the collector is down. UDP drivers like gelf over udp:// send and forget: the container starts and runs normally, and lines are lost on the wire without any error. Drivers that open a connection fail at start instead:
Without a listener on port 24224, docker run fails with exit status 125 and the container is left in Created: the logging driver is initialised as part of starting the task, and a refused connection aborts the start. That is loud, which is good, but it means a collector outage blocks every deploy on the node. fluentd-async=true lets the container start and buffers in the background while the driver reconnects. For the UDP drivers the only warning is the one you build: alert when the collector stops receiving from a host.
Blocking and non-blocking delivery
By default delivery is mode=blocking: when the driver cannot keep up, the application's write to stdout blocks until it can. That protects every line and can stall a request path when a remote collector is slow. mode=non-blocking puts a ring buffer in between (max-buffer-size, default 1 MB); the application never waits, and while the buffer is full, new lines are dropped without any error or metric. Keep the default blocking for audit-relevant logs and opt in to non-blocking per container for high-volume services where latency matters more than completeness: --log-opt mode=non-blocking --log-opt max-buffer-size=4m, or logging: options: in Compose. Both options also work as daemon defaults, but a host-wide non-blocking default silently allows drops for every container, audit logs included.
Changing the default for the whole host
This part runs on the sec VM. It follows the change procedure from "Configuring the daemon safely": back up the current file (or record an empty {} when there is none), merge the new keys into the backup with jq so settings already on the host survive, validate, restart, verify, and roll back by restoring the backup. The lesson's daemon.json makes local the default and sets rotation explicitly; delivery stays blocking:
{"log-driver": "local","log-opts": {"max-size": "10m","max-file": "3"}}
jq -s '.[0] * .[1]' merges the two objects recursively, with the lesson's keys winning; anything else in the original file (on VMs created with --mtu, the kit's MTU settings) is kept. lab-before was created before the change. Its restart policy brought it back after the daemon restart, still on json-file with no options, because a container's logging configuration is fixed when it is created. lab-after carries every default from the file. The only way to move an existing container to the new default is to recreate it:
Restoring the backup put the host back on json-file. After a default change on a real host, use docker inspect as above to find containers whose LogConfig.Type is still the old driver, and recreate them in a planned window.
What the application owes the driver
Drivers only see stdout and stderr. A process that writes to /var/log/app.log inside the container produces nothing for any driver, fills the container's writable layer instead, and disappears with the container. Runtimes that buffer output when they are not attached to a terminal look silent under docker logs until a buffer flushes; for Python, set PYTHONUNBUFFERED=1. One event per line, ideally as JSON, keeps collectors from splitting stack traces into dozens of records. Keep secrets out of log lines: they end up in the json-file, the journal, the dual-logging cache and the collector.
Clean up the main VM:
--log-driver fluentd, the other gelf with a udp:// address, both with default options. What do you see when you start them?"log-driver": "local" in daemon.json and restart dockerd. A week later docker inspect shows a long-running container with a restart policy is still on json-file. Why?Try this
Work through “What the application owes the driver” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from logging and log drivers, keep “What the application owes the driver”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.