Configuring the daemon safely

daemon.json, what to leave alone, and how to change it without surprises.

Intermediate20 min · lesson 8 of 24
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (3 files, 1 KB): daemonconfig.tar.gz. The lab VM shares no folders with your computer, so fetch them inside the VM: cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-hard/daemonconfig.tar.gz && tar -xzf daemonconfig.tar.gz, which creates ~/lab/daemonconfig/. SHA-256: bfd874b153cc183c0ea713cf50a81aca4f476208bc866a9405846ed08f57b424
Watch out
Run this only in the SecOpsLog disposable lab VM (secopslog-docker-sec), never on a host anyone depends on. The lab edits /etc/docker/daemon.json, restarts dockerd many times and deliberately breaks it once. If that VM does not exist yet, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, and open a shell in it with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec with Lima). To get back to a clean state at any point, recreate it with ./setup/create-lab.sh --profile sec --recreate. On this VM the ubuntu account is not in the docker group, so every docker command uses sudo. The lesson restarts dockerd more often than systemd allows by default (three starts a minute); when a restart fails with start request repeated too quickly, run sudo systemctl reset-failed docker and repeat it, or wait a minute.

Here is what an ordinary systemctl restart docker does to a running container on a default install:

ubuntu@secopslog-docker-sec:~/lab/daemonconfig · Docker 29.8.2
$ sudo cat /etc/docker/daemon.json 2>/dev/null || echo 'no daemon.json' sudo docker info --format 'store={{.Driver}} {{json .DriverStatus}} log={{.LoggingDriver}} live-restore={{.LiveRestoreEnabled}} root={{.DockerRootDir}}'
no daemon.json store=overlayfs [["driver-type","io.containerd.snapshotter.v1"]] log=json-file live-restore=false root=/var/lib/docker
$ sudo docker pull -q nginx:1.30-alpine sudo docker run -d --name lab-web nginx:1.30-alpine
docker.io/library/nginx:1.30-alpine 86ff9edf517a8490f6ba975e99f2be0f6edfb32d826bb1e0ee489025d95d8d69
$ sudo systemctl restart docker sudo docker ps -a --filter name=lab-web --format '{{.Names}}: {{.Status}}'
lab-web: Exited (0) Less than a second ago

There is no daemon.json: a fresh Docker 29 install runs entirely on built-in defaults, which the docker info line summarises (containerd image store, json-file logging, live-restore off, data under /var/lib/docker). The restart then stopped lab-web. With live-restore off, dockerd stops every container when it shuts down, so a daemon restart is a restart of every workload on the host. Most daemon changes need a restart, and this lesson is about making those changes without that kind of surprise: what belongs in daemon.json, how to change it safely, and which settings behave differently on Docker 29 than older guides say. Do not assume the file is missing on your own VMs: when the lab kit creates a VM on a link with an MTU below 1500 (typical behind a VPN), it writes a daemon.json with mtu settings, and the procedure below keeps them.

ubuntu@secopslog-docker-sec:~/lab/daemonconfig · Docker 29.8.2
$ sudo docker start lab-web sudo docker ps --filter name=lab-web --format '{{.Names}}: {{.Status}}'
lab-web lab-web: Up Less than a second

What goes in daemon.json

/etc/docker/daemon.json sets host-wide defaults for dockerd. Container flags still override many of them, but every container that does not say otherwise inherits them. A small file is a good file. This one sets the three things most production hosts want:

daemon.json
{
"live-restore": true,
"log-driver": "json-file",
"log-opts": {
"max-size": "10m",
"max-file": "3"
},
"default-address-pools": [
{ "base": "10.200.0.0/16", "size": 24 }
]
}

Other keys you will meet: registry-mirrors for a pull-through cache, features for feature switches (on Docker 29 the one that matters is containerd-snapshotter, shown below), firewall-backend for the iptables or nftables choice that "Publishing ports and the packet path" covers, and userns-remap, which turns on user-namespace remapping and, on Docker 29, disables the containerd image store; "User-namespace remapping" in Advanced container security explains that trade. Security defaults live here too: no-new-privileges for every new container, and seccomp-profile, a path to the seccomp profile used instead of the built-in one ("seccomp: filter system calls" in Advanced container security shows how to build one from the default).

Validate before you touch the live file

dockerd --validate --config-file parses a file and checks every key against the daemon's options without starting anything. Run it on the new file where you edited it, before it replaces the live one:

ubuntu@secopslog-docker-sec:~/lab/daemonconfig · Docker 29.8.2
$ sudo dockerd --validate --config-file ./daemon.json
configuration OK
$ sudo dockerd --validate --config-file ./typo-comma.json
unable to configure the Docker daemon with file ./typo-comma.json: invalid JSON: invalid character '}' looking for beginning of object key string
$ sudo dockerd --validate --config-file ./typo-key.json
unable to configure the Docker daemon with file ./typo-key.json: the following directives don't match any configuration option: max-size
typo-comma.json
{
"live-restore": true,
"log-driver": "json-file",
}
typo-key.json
{
"live-restore": true,
"log-opt": { "max-size": "10m" }
}

The trailing comma is a JSON syntax error, reported with the character the parser choked on. The second file is valid JSON with a misspelled key: log-opt instead of log-opts. The validator flattens unknown objects, so the message names the nested key max-size rather than the misspelled parent; when a key you know is valid is reported as unknown, check the spelling of the object around it. Both exit with status 1, which makes dockerd --validate easy to use as a gate in Ansible, Puppet or a CI check of the file.

Back up, merge, validate, install, reload, restart, verify

Never write a new daemon.json over the old one blindly. Something may already be in it (the kit's MTU settings, a registry mirror, a log default someone relies on), and a rollback needs the exact previous file. The procedure:

Other lessons in this course that touch daemon.json on the sec VM use the same steps, sometimes writing the merged file straight to /etc/docker/daemon.json and validating it there before the restart.

ubuntu@secopslog-docker-sec:~/lab/daemonconfig · Docker 29.8.2
$ sudo cp -a /etc/docker/daemon.json /etc/docker/daemon.json.bak 2>/dev/null || echo '{}' | sudo tee /etc/docker/daemon.json.bak >/dev/null sudo cat /etc/docker/daemon.json.bak
{}
$ jq -s '.[0] * .[1]' /etc/docker/daemon.json.bak daemon.json > merged.json cat merged.json sudo dockerd --validate --config-file ./merged.json
{ "live-restore": true, "log-driver": "json-file", "log-opts": { "max-size": "10m", "max-file": "3" }, "default-address-pools": [ { "base": "10.200.0.0/16", "size": 24 } ] } configuration OK
$ sudo install -m 0644 -o root -g root ./merged.json /etc/docker/daemon.json
$ sudo systemctl reload docker sudo docker info --format 'store={{.Driver}} {{json .DriverStatus}} log={{.LoggingDriver}} live-restore={{.LiveRestoreEnabled}} root={{.DockerRootDir}}'
store=overlayfs [["driver-type","io.containerd.snapshotter.v1"]] log=json-file live-restore=true root=/var/lib/docker

The backup holds {} because this VM had no file. jq -s reads both files into an array, and * merges the two objects recursively, so nested objects such as log-opts combine instead of replacing each other, and on a key present in both, the lesson's file wins. The merged result is validated where it was written, before it replaces the live file.

systemctl reload docker sends SIGHUP to dockerd, which re-reads daemon.json and applies the subset of options documented as reloadable, without stopping anything. live-restore is one of them, and docker info now reports live-restore=true. Log options are not on the reloadable list, so they wait for the restart (the lab-new container below shows them arriving). That order matters on a host that has never had live-restore: reload first so live-restore is active, then restart for the remaining keys.

ubuntu@secopslog-docker-sec:~/lab/daemonconfig · Docker 29.8.2
$ systemctl show docker -p MainPID pgrep -f 'shim-runc-v2 .*-id '$(sudo docker inspect -f '{{.Id}}' lab-web)
MainPID=2295 2561
$ sudo systemctl restart docker systemctl show docker -p MainPID pgrep -f 'shim-runc-v2 .*-id '$(sudo docker inspect -f '{{.Id}}' lab-web) sudo docker ps --filter name=lab-web --format '{{.Names}}: {{.Status}}'
MainPID=2978 2561 lab-web: Up 1 second
$ sudo docker info --format 'store={{.Driver}} {{json .DriverStatus}} log={{.LoggingDriver}} live-restore={{.LiveRestoreEnabled}} root={{.DockerRootDir}}'
store=overlayfs [["driver-type","io.containerd.snapshotter.v1"]] log=json-file live-restore=true root=/var/lib/docker

MainPID changed, so dockerd really restarted. The PID of lab-web's shim stayed the same, which proves the container process was never stopped; the short Up time in docker ps counts from the docker start a few seconds earlier, so it says little either way. The shim, which "Engine architecture on Docker 29" showed is the container's parent, kept the process alive while the daemon was gone, and the new dockerd reconnected to it. Live-restore covers daemon restarts and patch-level upgrades. It does not survive a host reboot, it does not apply to Swarm services, and the docs warn that it may fail when settings such as the bridge IP or the storage driver change between restarts. A long outage can also fill the container's output pipe and block the application's writes until dockerd returns.

Defaults reach new containers only

ubuntu@secopslog-docker-sec:~/lab/daemonconfig · Docker 29.8.2
$ sudo docker run -d --name lab-new nginx:1.30-alpine
e1ec42533e61675b7917d1b7d9f166af06a82445b975442dece32a1b0233920c
$ sudo docker inspect -f '{{.Name}} {{json .HostConfig.LogConfig}}' lab-web lab-new
/lab-web {"Type":"json-file","Config":{}} /lab-new {"Type":"json-file","Config":{"max-file":"3","max-size":"10m"}}

lab-new, created after the restart, carries max-file: 3 and max-size: 10m in its own HostConfig.LogConfig. lab-web was created before daemon.json existed and still has an empty config, meaning json-file with no rotation. The daemon copies defaults into a container's configuration when the container is created, so changing a default never touches existing containers; you have to recreate them (with Compose, docker compose up -d --force-recreate). Keep a list of what you recreated, or the next disk-full incident will be an old container nobody rolled.

ubuntu@secopslog-docker-sec:~/lab/daemonconfig · Docker 29.8.2
$ sudo docker network create lab-net-a sudo docker network create lab-net-b sudo docker network inspect -f '{{.Name}} {{range .IPAM.Config}}{{.Subnet}}{{end}}' lab-net-a lab-net-b ip -4 -br addr show docker0
f4322d873ddc6460a58be5f7eb68408384341aea126ab5ce9d659a116d4f51e2 37bcf94be417e92d1efd7bb0a289af8b1b34ca4b5b730fe49e5ee47fc4b508dc lab-net-a 10.200.0.0/24 lab-net-b 10.200.1.0/24 docker0 UP 172.17.0.1/16

Both user-defined networks came out of 10.200.0.0/16, one /24 each, starting at 10.200.0.0/24. docker0, the default bridge, kept 172.17.0.1/16, and that is a consequence of live-restore, not a rule. Without bip, dockerd reuses the address the bridge already has, and with live-restore on it leaves the existing bridge alone. With live-restore off, dockerd deletes and re-creates docker0 at startup, and the new one gets the first free subnet from the pools (10.200.0.1/24 with this file). The default bridge's address is also remembered in Docker's network store: remove the pools later and docker0 keeps the pool address. To decide the default bridge's range yourself, set bip (for example "bip": "10.250.0.1/24"), which wins in every case. Pick the pool after checking ip route on the host and the ranges your VPN and cloud networks use. Address clashes do not produce errors; traffic to the clashing network is routed into the container bridge and disappears.

A valid file that stops dockerd

A common way to expose the API on another socket or on TCP is a hosts key in daemon.json. On a systemd host this breaks the daemon, and the validator does not catch it:

ubuntu@secopslog-docker-sec:~/lab/daemonconfig · Docker 29.8.2
$ jq '. + {"hosts": ["unix:///var/run/docker.sock"]}' merged.json > hosts.json sudo dockerd --validate --config-file ./hosts.json
configuration OK
$ sudo install -m 0644 ./hosts.json /etc/docker/daemon.json sudo systemctl restart docker
Job for docker.service failed because the control process exited with error code. See "systemctl status docker.service" and "journalctl -xeu docker.service" for details.
$ systemctl is-active docker sudo journalctl -u docker -n 30 --no-pager -o cat | grep -m1 -i 'directives' sudo ctr -n moby tasks ls
activating unable to configure the Docker daemon with file /etc/docker/daemon.json: the following directives are specified both as a flag and in the configuration file: hosts: (from flag: [fd://], from file: [unix:///var/run/docker.sock]) TASK PID STATUS 86ff9edf517a8490f6ba975e99f2be0f6edfb32d826bb1e0ee489025d95d8d69 2584 RUNNING e1ec42533e61675b7917d1b7d9f166af06a82445b975442dece32a1b0233920c 3302 RUNNING

--validate saw only the file and said configuration OK. The real start also has the -H fd:// flag from the unit's ExecStart, and dockerd refuses to start when an option is set both as a flag and in the file; the journal names the clash exactly. activating means systemd is retrying under Restart=always, and it gives up after three starts within a minute. Meanwhile ctr -n moby tasks ls shows both containers still running: with live-restore on, a dead daemon is an API outage, not a workload outage. The fix is to reinstall the last known-good file, here merged.json:

ubuntu@secopslog-docker-sec:~/lab/daemonconfig · Docker 29.8.2
$ sudo install -m 0644 ./merged.json /etc/docker/daemon.json sudo systemctl reset-failed docker sudo systemctl restart docker systemctl is-active docker sudo docker ps --filter name=lab- --format '{{.Names}}: {{.Status}}'
active lab-new: Up 1 second lab-web: Up 2 seconds

reset-failed clears systemd's start-limit counter so the next start is attempted. If you really need another listener, add it with a systemd drop-in that overrides ExecStart, and never on plain TCP without TLS client certificates; "The Docker socket and daemon hardening" in Advanced container security covers that.

data-root does not move images on Docker 29

data-root used to move everything Docker stored. With the containerd image store, image content and container snapshots belong to containerd and live under containerd's own root. The containers are removed first because changing the data directory under running containers is one of the cases where live-restore cannot reconnect.

ubuntu@secopslog-docker-sec:~/lab/daemonconfig · Docker 29.8.2
$ sudo docker rm -f lab-web lab-new
lab-web lab-new
$ jq '. + {"data-root": "/srv/docker"}' merged.json > data-root.json sudo dockerd --validate --config-file ./data-root.json sudo install -m 0644 ./data-root.json /etc/docker/daemon.json sudo systemctl restart docker sudo docker info --format 'store={{.Driver}} {{json .DriverStatus}} log={{.LoggingDriver}} live-restore={{.LiveRestoreEnabled}} root={{.DockerRootDir}}'
configuration OK store=overlayfs [["driver-type","io.containerd.snapshotter.v1"]] log=json-file live-restore=true root=/srv/docker
$ sudo docker image ls nginx sudo docker network ls --filter name=lab-
IMAGE ID DISK USAGE CONTENT SIZE EXTRA nginx:1.30-alpine 0985e772fb9f 92.9MB 26.9MB NETWORK ID NAME DRIVER SCOPE
$ sudo du -sh /srv/docker /var/lib/docker /var/lib/containerd
216K /srv/docker 272K /var/lib/docker 5.7G /var/lib/containerd

docker info reports root=/srv/docker, yet docker image ls still lists nginx and du shows where image data lives: everything the VM has pulled is in /var/lib/containerd, against a few hundred kilobytes in each Docker directory. The lab- networks are gone because networks, volumes and container records are dockerd's data and stayed in the old /var/lib/docker. To move images to another disk, change containerd's root in /etc/containerd/config.toml as well, with containers stopped, both daemons stopped and the data copied; "Where Docker keeps data" walks through it. Moving back restores the networks:

ubuntu@secopslog-docker-sec:~/lab/daemonconfig · Docker 29.8.2
$ sudo install -m 0644 ./merged.json /etc/docker/daemon.json sudo systemctl restart docker sudo docker network ls --filter name=lab-
NETWORK ID NAME DRIVER SCOPE f4322d873ddc lab-net-a bridge local 37bcf94be417 lab-net-b bridge local

storage-driver: overlay2 switches the image store

Older guides put "storage-driver": "overlay2" in every daemon.json. On a fresh Docker 29 host that line does not describe the current setup, it replaces it with the legacy graph-driver image store:

ubuntu@secopslog-docker-sec:~/lab/daemonconfig · Docker 29.8.2
$ jq '. + {"storage-driver": "overlay2"}' merged.json > storage-driver.json sudo install -m 0644 ./storage-driver.json /etc/docker/daemon.json sudo systemctl restart docker sudo docker info --format 'store={{.Driver}} {{json .DriverStatus}} log={{.LoggingDriver}} live-restore={{.LiveRestoreEnabled}} root={{.DockerRootDir}}'
store=overlay2 [["Backing Filesystem","extfs"],["Supports d_type","true"],["Using metacopy","false"],["Native Overlay Diff","true"],["userxattr","false"]] log=json-file live-restore=true root=/var/lib/docker
$ sudo docker image ls nginx sudo ctr -n moby images ls -q | grep nginx
IMAGE ID DISK USAGE CONTENT SIZE EXTRA docker.io/library/nginx:1.30-alpine

docker info now reports overlay2 with the classic graph-driver details, and docker image ls finds no nginx, while containerd still holds it. Nothing was deleted; each store hides the other's images. The surprise comes when you remove the line again:

ubuntu@secopslog-docker-sec:~/lab/daemonconfig · Docker 29.8.2
$ sudo install -m 0644 ./merged.json /etc/docker/daemon.json sudo systemctl restart docker sudo docker info --format '{{.Driver}}' sudo docker image ls nginx
overlay2 IMAGE ID DISK USAGE CONTENT SIZE EXTRA
$ sudo ls /var/lib/docker /var/lib/docker/image
/var/lib/docker: buildkit containers engine-id image network overlay2 plugins rootfs runtimes swarm tmp volumes /var/lib/docker/image: identity-cache.db overlay2

The daemon stays on overlay2. Once /var/lib/docker/overlay2 and /var/lib/docker/image/overlay2 exist, dockerd treats the host like one upgraded from an older engine and keeps the legacy store, which is how upgraded hosts avoid losing their images. The way back is to ask for the containerd store explicitly:

ubuntu@secopslog-docker-sec:~/lab/daemonconfig · Docker 29.8.2
$ jq '. + {"features": {"containerd-snapshotter": true}}' merged.json > snapshotter.json sudo dockerd --validate --config-file ./snapshotter.json sudo install -m 0644 ./snapshotter.json /etc/docker/daemon.json sudo systemctl restart docker sudo docker info --format '{{.Driver}} {{json .DriverStatus}}' sudo docker image ls nginx
configuration OK overlayfs [["driver-type","io.containerd.snapshotter.v1"]] IMAGE ID DISK USAGE CONTENT SIZE EXTRA nginx:1.30-alpine 0985e772fb9f 92.9MB 26.9MB

features.containerd-snapshotter: true selects the containerd store regardless of leftover graph-driver state, and nginx is visible again. The same flag set to false is how you would deliberately stay on the legacy store, and set to true on a host upgraded from Docker 28 it is the migration switch. Images and containers in the old store are hidden after it: docker save the images you need before switching and docker load them after, or pull and rebuild them, then recreate the containers. On a fresh Docker 29 host, leave both storage-driver and the feature flag out.

Roll back

Rolling back means restoring the backup. Here that is the {} recorded at the start, which restores the defaults for everything except the image store, which stays on the legacy store for the reason above:

ubuntu@secopslog-docker-sec:~/lab/daemonconfig · Docker 29.8.2
$ sudo docker network rm lab-net-a lab-net-b sudo cp /etc/docker/daemon.json.bak /etc/docker/daemon.json sudo systemctl restart docker sudo cat /etc/docker/daemon.json sudo docker info --format 'store={{.Driver}} {{json .DriverStatus}} log={{.LoggingDriver}} live-restore={{.LiveRestoreEnabled}} root={{.DockerRootDir}}'
lab-net-a lab-net-b {} store=overlay2 [["Backing Filesystem","extfs"],["Supports d_type","true"],["Using metacopy","false"],["Native Overlay Diff","true"],["userxattr","false"]] log=json-file live-restore=false root=/var/lib/docker
$ sudo systemctl stop docker docker.socket sudo rm -rf /var/lib/docker/overlay2 /var/lib/docker/image/overlay2 /srv/docker sudo systemctl start docker sudo docker info --format 'store={{.Driver}} {{json .DriverStatus}} log={{.LoggingDriver}} live-restore={{.LiveRestoreEnabled}} root={{.DockerRootDir}}' sudo docker image ls nginx rm -f merged.json hosts.json data-root.json storage-driver.json snapshotter.json
store=overlayfs [["driver-type","io.containerd.snapshotter.v1"]] log=json-file live-restore=false root=/var/lib/docker IMAGE ID DISK USAGE CONTENT SIZE EXTRA nginx:1.30-alpine 0985e772fb9f 92.9MB 26.9MB
# Lab only: deletes the empty legacy graph-driver directories. Never do this on a host whose legacy store holds images.

Deleting the two overlay2 directories is safe here only because the legacy store never held an image on this VM. On a real host those directories are where an upgraded engine keeps its images, so do not copy that step outside the lab.

Keys to leave alone

Keep daemon.json in the same repository as the rest of the host build, run dockerd --validate on it in CI, and roll changes to one canary host before the rest.

Quick check
01You add a hosts key to daemon.json. dockerd --validate --config-file prints configuration OK, but systemctl restart docker fails and the journal says the directives are specified both as a flag and in the configuration file. Why did validation pass?
Incorrect — It does check key names; the typo-key file in the lesson fails on an unknown directive.
Incorrect — Permissions are not the issue; the validator never looks at the unit at all.
Correct — The clash exists only when the file and the unit's flags meet in a real start, so the file on its own is valid.
Incorrect — Reloadability has nothing to do with validation, and hosts is not reloadable.
02A busy host has never had daemon.json. You must turn on live-restore and set log rotation, and no container may stop. Which order works?
Correct — live-restore is reloadable, so it is already active when the restart happens, and the shims keep the containers running.
Incorrect — live-restore is not active yet during that restart, so dockerd stops every container on the way down.
Incorrect — The containers would come back, but they still stop and restart, which is the outage you wanted to avoid.
Incorrect — dockerd reads the file at start and on reload only; nothing watches it.
03On a fresh Docker 29 host you set "data-root": "/data/docker" to move Docker off the small root disk. After the restart, pulls still fill the root disk. What did you miss?
Incorrect — Setting storage-driver: overlay2 would switch to the legacy store and hide your images; it is not part of moving data-root.
Incorrect — The lesson's lab shows docker info reporting the new root right after a daemon restart.
Incorrect — There is no such staging step; the image data simply lives elsewhere.
Correct — With the containerd image store, data-root moves dockerd's data only; change containerd's root in /etc/containerd/config.toml too.

Try this

Work through “Keys to leave alone” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from configuring the daemon safely, keep “Keys to leave alone”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related