Configuring the daemon safely
daemon.json, what to leave alone, and how to change it without surprises.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-hard/daemonconfig.tar.gz && tar -xzf daemonconfig.tar.gz, which creates ~/lab/daemonconfig/. SHA-256: bfd874b153cc183c0ea713cf50a81aca4f476208bc866a9405846ed08f57b424./setup/create-lab.sh --profile sec, and open a shell in it with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec with Lima). To get back to a clean state at any point, recreate it with ./setup/create-lab.sh --profile sec --recreate. On this VM the ubuntu account is not in the docker group, so every docker command uses sudo. The lesson restarts dockerd more often than systemd allows by default (three starts a minute); when a restart fails with start request repeated too quickly, run sudo systemctl reset-failed docker and repeat it, or wait a minute.Here is what an ordinary systemctl restart docker does to a running container on a default install:
There is no daemon.json: a fresh Docker 29 install runs entirely on built-in defaults, which the docker info line summarises (containerd image store, json-file logging, live-restore off, data under /var/lib/docker). The restart then stopped lab-web. With live-restore off, dockerd stops every container when it shuts down, so a daemon restart is a restart of every workload on the host. Most daemon changes need a restart, and this lesson is about making those changes without that kind of surprise: what belongs in daemon.json, how to change it safely, and which settings behave differently on Docker 29 than older guides say. Do not assume the file is missing on your own VMs: when the lab kit creates a VM on a link with an MTU below 1500 (typical behind a VPN), it writes a daemon.json with mtu settings, and the procedure below keeps them.
What goes in daemon.json
/etc/docker/daemon.json sets host-wide defaults for dockerd. Container flags still override many of them, but every container that does not say otherwise inherits them. A small file is a good file. This one sets the three things most production hosts want:
{"live-restore": true,"log-driver": "json-file","log-opts": {"max-size": "10m","max-file": "3"},"default-address-pools": [{ "base": "10.200.0.0/16", "size": 24 }]}
Other keys you will meet: registry-mirrors for a pull-through cache, features for feature switches (on Docker 29 the one that matters is containerd-snapshotter, shown below), firewall-backend for the iptables or nftables choice that "Publishing ports and the packet path" covers, and userns-remap, which turns on user-namespace remapping and, on Docker 29, disables the containerd image store; "User-namespace remapping" in Advanced container security explains that trade. Security defaults live here too: no-new-privileges for every new container, and seccomp-profile, a path to the seccomp profile used instead of the built-in one ("seccomp: filter system calls" in Advanced container security shows how to build one from the default).
Validate before you touch the live file
dockerd --validate --config-file parses a file and checks every key against the daemon's options without starting anything. Run it on the new file where you edited it, before it replaces the live one:
{"live-restore": true,"log-driver": "json-file",}
{"live-restore": true,"log-opt": { "max-size": "10m" }}
The trailing comma is a JSON syntax error, reported with the character the parser choked on. The second file is valid JSON with a misspelled key: log-opt instead of log-opts. The validator flattens unknown objects, so the message names the nested key max-size rather than the misspelled parent; when a key you know is valid is reported as unknown, check the spelling of the object around it. Both exit with status 1, which makes dockerd --validate easy to use as a gate in Ansible, Puppet or a CI check of the file.
Back up, merge, validate, install, reload, restart, verify
Never write a new daemon.json over the old one blindly. Something may already be in it (the kit's MTU settings, a registry mirror, a log default someone relies on), and a rollback needs the exact previous file. The procedure:
Other lessons in this course that touch daemon.json on the sec VM use the same steps, sometimes writing the merged file straight to /etc/docker/daemon.json and validating it there before the restart.
The backup holds {} because this VM had no file. jq -s reads both files into an array, and * merges the two objects recursively, so nested objects such as log-opts combine instead of replacing each other, and on a key present in both, the lesson's file wins. The merged result is validated where it was written, before it replaces the live file.
systemctl reload docker sends SIGHUP to dockerd, which re-reads daemon.json and applies the subset of options documented as reloadable, without stopping anything. live-restore is one of them, and docker info now reports live-restore=true. Log options are not on the reloadable list, so they wait for the restart (the lab-new container below shows them arriving). That order matters on a host that has never had live-restore: reload first so live-restore is active, then restart for the remaining keys.
MainPID changed, so dockerd really restarted. The PID of lab-web's shim stayed the same, which proves the container process was never stopped; the short Up time in docker ps counts from the docker start a few seconds earlier, so it says little either way. The shim, which "Engine architecture on Docker 29" showed is the container's parent, kept the process alive while the daemon was gone, and the new dockerd reconnected to it. Live-restore covers daemon restarts and patch-level upgrades. It does not survive a host reboot, it does not apply to Swarm services, and the docs warn that it may fail when settings such as the bridge IP or the storage driver change between restarts. A long outage can also fill the container's output pipe and block the application's writes until dockerd returns.
Defaults reach new containers only
lab-new, created after the restart, carries max-file: 3 and max-size: 10m in its own HostConfig.LogConfig. lab-web was created before daemon.json existed and still has an empty config, meaning json-file with no rotation. The daemon copies defaults into a container's configuration when the container is created, so changing a default never touches existing containers; you have to recreate them (with Compose, docker compose up -d --force-recreate). Keep a list of what you recreated, or the next disk-full incident will be an old container nobody rolled.
Both user-defined networks came out of 10.200.0.0/16, one /24 each, starting at 10.200.0.0/24. docker0, the default bridge, kept 172.17.0.1/16, and that is a consequence of live-restore, not a rule. Without bip, dockerd reuses the address the bridge already has, and with live-restore on it leaves the existing bridge alone. With live-restore off, dockerd deletes and re-creates docker0 at startup, and the new one gets the first free subnet from the pools (10.200.0.1/24 with this file). The default bridge's address is also remembered in Docker's network store: remove the pools later and docker0 keeps the pool address. To decide the default bridge's range yourself, set bip (for example "bip": "10.250.0.1/24"), which wins in every case. Pick the pool after checking ip route on the host and the ranges your VPN and cloud networks use. Address clashes do not produce errors; traffic to the clashing network is routed into the container bridge and disappears.
A valid file that stops dockerd
A common way to expose the API on another socket or on TCP is a hosts key in daemon.json. On a systemd host this breaks the daemon, and the validator does not catch it:
--validate saw only the file and said configuration OK. The real start also has the -H fd:// flag from the unit's ExecStart, and dockerd refuses to start when an option is set both as a flag and in the file; the journal names the clash exactly. activating means systemd is retrying under Restart=always, and it gives up after three starts within a minute. Meanwhile ctr -n moby tasks ls shows both containers still running: with live-restore on, a dead daemon is an API outage, not a workload outage. The fix is to reinstall the last known-good file, here merged.json:
reset-failed clears systemd's start-limit counter so the next start is attempted. If you really need another listener, add it with a systemd drop-in that overrides ExecStart, and never on plain TCP without TLS client certificates; "The Docker socket and daemon hardening" in Advanced container security covers that.
data-root does not move images on Docker 29
data-root used to move everything Docker stored. With the containerd image store, image content and container snapshots belong to containerd and live under containerd's own root. The containers are removed first because changing the data directory under running containers is one of the cases where live-restore cannot reconnect.
docker info reports root=/srv/docker, yet docker image ls still lists nginx and du shows where image data lives: everything the VM has pulled is in /var/lib/containerd, against a few hundred kilobytes in each Docker directory. The lab- networks are gone because networks, volumes and container records are dockerd's data and stayed in the old /var/lib/docker. To move images to another disk, change containerd's root in /etc/containerd/config.toml as well, with containers stopped, both daemons stopped and the data copied; "Where Docker keeps data" walks through it. Moving back restores the networks:
storage-driver: overlay2 switches the image store
Older guides put "storage-driver": "overlay2" in every daemon.json. On a fresh Docker 29 host that line does not describe the current setup, it replaces it with the legacy graph-driver image store:
docker info now reports overlay2 with the classic graph-driver details, and docker image ls finds no nginx, while containerd still holds it. Nothing was deleted; each store hides the other's images. The surprise comes when you remove the line again:
The daemon stays on overlay2. Once /var/lib/docker/overlay2 and /var/lib/docker/image/overlay2 exist, dockerd treats the host like one upgraded from an older engine and keeps the legacy store, which is how upgraded hosts avoid losing their images. The way back is to ask for the containerd store explicitly:
features.containerd-snapshotter: true selects the containerd store regardless of leftover graph-driver state, and nginx is visible again. The same flag set to false is how you would deliberately stay on the legacy store, and set to true on a host upgraded from Docker 28 it is the migration switch. Images and containers in the old store are hidden after it: docker save the images you need before switching and docker load them after, or pull and rebuild them, then recreate the containers. On a fresh Docker 29 host, leave both storage-driver and the feature flag out.
Roll back
Rolling back means restoring the backup. Here that is the {} recorded at the start, which restores the defaults for everything except the image store, which stays on the legacy store for the reason above:
Deleting the two overlay2 directories is safe here only because the legacy store never held an image on this VM. On a real host those directories are where an upgraded engine keeps its images, so do not copy that step outside the lab.
Keys to leave alone
Keep daemon.json in the same repository as the rest of the host build, run dockerd --validate on it in CI, and roll changes to one canary host before the rest.
hosts key to daemon.json. dockerd --validate --config-file prints configuration OK, but systemctl restart docker fails and the journal says the directives are specified both as a flag and in the configuration file. Why did validation pass?"data-root": "/data/docker" to move Docker off the small root disk. After the restart, pulls still fill the root disk. What did you miss?Try this
Work through “Keys to leave alone” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from configuring the daemon safely, keep “Keys to leave alone”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.