Networks, drivers and DNS

Bridges, veth pairs, the embedded DNS server, internal networks, ipvlan and legacy links.

Intermediate15 min · lesson 12 of 24

Use the main lab VM (secopslog-docker) for this lesson. Two settings decide how a container is networked: the driver of each network it joins, and which networks those are. Start with what this daemon can build:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker info --format '{{.Plugins.Network}}' docker network ls --filter type=builtin
[bridge host ipvlan macvlan null overlay] NETWORK ID NAME DRIVER SCOPE a61cea6118e5 bridge bridge local 3052775d188b host host local 1b4aa92363a3 none null local

The first line lists the network drivers compiled into this Docker 29.8.2 daemon. The table lists the three networks every install creates: bridge (the default bridge, the one a container joins when you pass no --network), host and none (driver null). "Container networking basics" (Docker for beginners) showed the everyday difference between the default bridge and a network you create. This lesson opens those networks up: what the kernel objects are, how name resolution is wired in, and what the other drivers give you.

What a bridge network is made of

Create a user-defined bridge network with a fixed subnet (fixed so the addresses below stay the same between runs; without --subnet Docker picks one from its address pools) and read back what Docker recorded:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker network create --subnet 10.89.1.0/24 lab-front
9c22ab120a56b6f3626db543d65af8a260645172b32e45dc8fd21e331b72d007
$ docker network inspect lab-front --format '{{.Driver}} ipv4={{.EnableIPv4}} ipv6={{.EnableIPv6}} internal={{.Internal}} {{json .IPAM.Config}}'
bridge ipv4=true ipv6=false internal=false [{"Subnet":"10.89.1.0/24","Gateway":"10.89.1.1"}]
$ docker run -d --name lab-web --network lab-front nginx:1.30-alpine
0bbe0b4429064189c272881f3a0bc0c80afc8eea29f938b40173b599b4387b28

ipv4=true ipv6=false is the default: IPv6 is opt-in per network, which matters in "Publishing ports and the packet path". internal=false means the network has a route to the outside. The gateway 10.89.1.1 is not a container; it is an address Docker puts on a Linux bridge interface in the host. Look for it, and for what is plugged into it:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ BR=br-$(docker network inspect -f '{{.Id}}' lab-front | cut -c1-12) ip -br addr show $BR ip link show master $BR
br-9c22ab120a56 UP 10.89.1.1/24 fe80::18ed:caff:fe22:55d/64 1209: vethf4537c0@if2: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue master br-9c22ab120a56 state UP mode DEFAULT group default link/ether 36:48:6e:35:85:a1 brd ff:ff:ff:ff:ff:ff link-netnsid 0

Docker names the bridge br- plus the first 12 characters of the network ID, so your name will differ. The gateway address sits on the bridge, and one interface is attached to it with master br-...: vethf4537c0, interface number 1209 on this host. A veth is a virtual cable with two ends, and the @if2 says the other end is interface 2 in another network namespace. The container's side looks like this:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ PID=$(docker inspect -f '{{.State.Pid}}' lab-web) sudo nsenter -n -t $PID ip -br addr sudo nsenter -n -t $PID ip route
lo UNKNOWN 127.0.0.1/8 ::1/128 eth0@if1209 UP 10.89.1.2/24 default via 10.89.1.1 dev eth0 10.89.1.0/24 dev eth0 proto kernel scope link src 10.89.1.2

nsenter -n -t PID runs a host command inside the network namespace of process PID, here the container's main process; it needs root, hence sudo, and works on images without any tools, which "Network troubleshooting" relies on. Inside, there is a loopback interface and eth0@if1209: the other end of the cable, pointing back at host interface 1209. The container's default route goes to the bridge address. Interface numbers, veth names and the bridge name change with every container and network; the pairing rule does not. "Namespaces from first principles" (Advanced container security) covers network namespaces next to the other seven kinds.

One container on a user-defined bridge
Host network namespace
enp0s1 (host NIC)
the host's own address; traffic to the outside leaves here
br-9c22ab120a56: 10.89.1.1/24
Linux bridge for lab-front, the containers' gateway
vethf4537c0 (ifindex 1209)
host end of the cable, attached to the bridge
Container network namespace (lab-web)
eth0@if1209: 10.89.1.2/24
container end of the same veth pair
default via 10.89.1.1
everything not on 10.89.1.0/24 goes to the bridge
127.0.0.11
embedded DNS address, answered by dockerd
Between the bridge and enp0s1 the host routes and filters; "Publishing ports and the packet path" follows a packet through those rules.

The embedded DNS server

Compare what a container sees as its resolver on the default bridge and on lab-front:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm alpine:3.22 cat /etc/resolv.conf
# Generated by Docker Engine. # This file can be edited; Docker Engine will not make further changes once it # has been modified. nameserver 192.168.2.1 search . # Based on host file: '/run/systemd/resolve/resolv.conf' (legacy) # Overrides: []
$ docker run --rm --network lab-front alpine:3.22 cat /etc/resolv.conf
# Generated by Docker Engine. # This file can be edited; Docker Engine will not make further changes once it # has been modified. nameserver 127.0.0.11 search . options edns0 trust-ad ndots:0 # Based on host file: '/etc/resolv.conf' (internal resolver) # ExtServers: [host(127.0.0.53)] # Overrides: [] # Option ndots from: internal

On the default bridge, Docker copies the host's upstream resolvers into the container (192.168.2.1 is this Multipass VM's gateway; yours will be whatever your host uses). Nothing in that file knows container names, which is why the default bridge has no name resolution between containers. On a user-defined network the only nameserver is 127.0.0.11, and the comment shows where unknown names go: ExtServers: [host(127.0.0.53)], the host's systemd-resolved, queried from the host's namespace. --dns on docker run changes those upstream servers, not the 127.0.0.11 line. Docker regenerates the file until you edit it inside the container, then leaves your edit alone, as the header says.

127.0.0.11 is a loopback address inside the container, yet no process in the container answers it. The answer is in the container's own NAT table:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ PID=$(docker inspect -f '{{.State.Pid}}' lab-web) sudo nsenter -n -t $PID iptables -t nat -S DOCKER_OUTPUT sudo nsenter -n -t $PID ss -lnup
-N DOCKER_OUTPUT -A DOCKER_OUTPUT -d 127.0.0.11/32 -p tcp -m tcp --dport 53 -j DNAT --to-destination 127.0.0.11:36849 -A DOCKER_OUTPUT -d 127.0.0.11/32 -p udp -m udp --dport 53 -j DNAT --to-destination 127.0.0.11:37240 State Recv-Q Send-Q Local Address:Port Peer Address:PortProcess UNCONN 0 0 127.0.0.11:37240 0.0.0.0:* users:(("dockerd",pid=17027,fd=45))

Docker installs rules inside each container's namespace that redirect port 53 on 127.0.0.11 to two random ports, and the process listening there is dockerd: the daemon opened sockets inside the container's namespace. The ports differ per container. Two consequences follow. Name resolution between containers depends on a running dockerd; with live-restore on, containers keep running while the daemon is down, but they cannot resolve each other's names until it is back. And a container with its own firewall tooling must not flush its NAT table, or DNS stops.

The server answers for container names, for the service names Compose registers, and for aliases. An alias is an extra name on one network; give the same alias to two containers and the server returns both addresses:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-api1 --network lab-front --network-alias api nginx:1.30-alpine docker run -d --name lab-api2 --network lab-front --network-alias api nginx:1.30-alpine
f79802ebe9a75f847055e349bdd672623d2137593b3573522b7654b2c4296bcf 8fc2b46ebf809c1f2d16263411f408b24d58bd23ac3b2e6bc2c40854293c526d
$ docker run --rm --network lab-front alpine:3.22 nslookup -type=a api
Server: 127.0.0.11 Address: 127.0.0.11:53 Non-authoritative answer: Name: api Address: 10.89.1.3 Name: api Address: 10.89.1.4

Both containers answer to api. The order of the records can change between queries, and most clients connect to the first, so this spreads load roughly at best and knows nothing about health: a stopped container drops out of the answers, a hung one does not. Clients also cache. The JVM keeps successful lookups for networkaddress.cache.ttl, which can be forever under a security manager, so after a container is recreated with a new address the application may keep dialling the old one until it reconnects or restarts. Swarm services use a virtual IP by default and DNS round robin only with endpoint_mode: dnsrr; "Overlay networks and the routing mesh" covers both.

Segmenting with networks

A container can reach and resolve only containers on networks it shares. Put a cache on a second network, created with --internal, and ask for it from lab-web:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker network create --internal --subnet 10.89.2.0/24 lab-back docker run -d --name lab-cache --network lab-back redis:8-alpine
d0fc071cb996eb3f54a5173eb9ab2f402c46af2abf85708bbe9c00c757502684 c74ce47823213298dda201c2ed305c51613ff39a211897b783877e3f6bcdbe28
$ docker exec lab-web nslookup lab-cache
Server: 127.0.0.11 Address: 127.0.0.11:53 ** server can't find lab-cache: SERVFAIL ** server can't find lab-cache: SERVFAIL

lab-cache is not on any network lab-web belongs to, so the embedded server does not know the name and passes it upstream, which cannot resolve it either: SERVFAIL here, NXDOMAIN with other upstream resolvers. Either way the cause is network membership, not the cache. Connect lab-web to the second network and the name and the port work:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker network connect lab-back lab-web docker exec lab-web nc -zv lab-cache 6379
lab-cache (10.89.2.2:6379) open
$ docker inspect -f '{{range $k, $v := .NetworkSettings.Networks}}{{$k}}={{$v.IPAddress}} {{end}}' lab-web
lab-back=10.89.2.3 lab-front=10.89.1.2

docker network connect added a second interface to the running container, and it now has one address per network. This is the usual two-tier shape: the web tier on a front network, the data tier on a back network that only the web tier joins. --internal on lab-back changes one more thing:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker exec lab-cache ip route
10.89.2.0/24 dev eth0 scope link src 10.89.2.2
$ docker exec lab-cache wget -T 3 -qO- http://example.com
wget: bad address 'example.com'

The cache has a route for its own subnet and no default route, and its DNS lookups for outside names fail, so it cannot reach anything beyond lab-back. That is an internal network: Docker sets up no gateway route and no masquerading for it. lab-web still reaches the internet through lab-front. Choosing which tiers get egress, and controlling it, is the subject of "Container network hardening" (Advanced container security).

none and host

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm --network none alpine:3.22 ip addr
1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc noqueue state UNKNOWN qlen 1000 link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00 inet 127.0.0.1/8 scope host lo valid_lft forever preferred_lft forever inet6 ::1/128 scope host valid_lft forever preferred_lft forever
$ sudo readlink /proc/1/ns/net docker run --rm --network host alpine:3.22 readlink /proc/self/ns/net docker run --rm alpine:3.22 readlink /proc/self/ns/net
net:[4026531833] net:[4026531833] net:[4026532639]

With --network none the namespace has loopback and nothing else. With --network host there is no separate namespace at all: /proc/self/ns/net in the container names the same namespace (4026531833) as PID 1 on the host, while an ordinary container gets its own. A host-network container binds host ports directly, -p is ignored, and nothing filters its traffic separately from the host's; treat it as a privilege, as "Host namespace sharing as attack surface" (Advanced container security) does.

ipvlan and macvlan

Both drivers attach containers to a parent interface instead of a bridge, so containers get addresses on the parent's network, with no NAT and no published ports. Without -o parent=, Docker creates a dummy parent, which is enough to see how ipvlan works without touching a real LAN:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker network create -d ipvlan --subnet 10.89.4.0/24 lab-ipvl docker run -d --name lab-iv1 --network lab-ipvl alpine:3.22 sleep 600 docker run -d --name lab-iv2 --network lab-ipvl alpine:3.22 sleep 600
0ed0452e5528dda959dcbab21bb77fd5bb2e7edf8faea586afa63f2eaa124a63 7020227348ecef788599ab67818eaa6bf79b34c1560bc786d3f4a4d85db70f69 8a1353b4726ea7084f55872b784ce7c3557bda07b25c8742d6d660b6bb18ead7
$ docker exec lab-iv1 ip link show eth0 docker exec lab-iv2 ip link show eth0 ip -br link show type dummy
1224: eth0@if1223: <BROADCAST,MULTICAST,UP,LOWER_UP,M-DOWN> mtu 1500 qdisc noqueue state UNKNOWN link/ether ea:ca:fc:02:11:45 brd ff:ff:ff:ff:ff:ff 1225: eth0@if1223: <BROADCAST,MULTICAST,UP,LOWER_UP,M-DOWN> mtu 1500 qdisc noqueue state UNKNOWN link/ether ea:ca:fc:02:11:45 brd ff:ff:ff:ff:ff:ff di-0ed0452e5528 UNKNOWN ea:ca:fc:02:11:45 <BROADCAST,NOARP,UP,LOWER_UP>
$ docker exec lab-iv1 ping -c 1 -W 2 lab-iv2
PING lab-iv2 (10.89.4.2): 56 data bytes 64 bytes from 10.89.4.2: seq=0 ttl=64 time=0.383 ms --- lab-iv2 ping statistics --- 1 packets transmitted, 1 packets received, 0% packet loss round-trip min/avg/max = 0.383/0.383/0.383 ms

Both containers have the same MAC address, the MAC of the dummy interface di-... that Docker created as the parent. That is ipvlan's defining property: every container shares the parent's MAC and is told apart by IP address, which suits switch ports and cloud networks that accept only one MAC per port. Containers on the same ipvlan network reach each other, here by name. With a real parent (-o parent=eth0, plus the LAN's --subnet and --gateway) they become reachable from the LAN; -o ipvlan_mode=l3 makes the host route between container subnets instead of bridging at layer 2. A network with a dummy parent is internal by design.

macvlan gives every container its own MAC address on the parent's network, so to the LAN each container looks like a separate machine. Three things to know before choosing it. The host cannot reach its own macvlan containers through the parent interface; the usual fix is a macvlan sub-interface on the host with a route to the container addresses. Many cloud and virtualised networks drop frames from MAC addresses they did not assign, so macvlan often fails there; ipvlan avoids that. Each container takes a real address from the LAN. Since Docker 29.0, macvlan and ipvlan L2 networks get no default gateway at all, IPv4 or IPv6, unless the IPAM configuration includes --gateway; give one explicitly when containers must route beyond the LAN. macvlan is described here and not run: the lab VM sits behind its hypervisor's NAT, so there is no LAN on which to show the result.

Legacy links

Old scripts and Compose files still use --link and links:. On Docker 29 a link on the default bridge does this:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-old -e LAB_SETTING=only-for-lab-old alpine:3.22 sleep 600 docker run --rm --link lab-old:old alpine:3.22 sh -c "grep old /etc/hosts; env | grep ^OLD_ || echo no OLD_ variables"
f3c13538a7607e82851ce032129263a09b7fe613567e0b2a16ba8a0a2ad52f15 WARNING: Links on the default bridge network are deprecated and will be removed in a future release. Use a custom network instead. 172.17.0.3 old f3c13538a760 lab-old no OLD_ variables
# Legacy flag, run here only to show what it still does on Docker 29.

Creating the container prints a deprecation warning (Docker 29.6 and later), the link writes one /etc/hosts entry with the target's address at start time, and no OLD_* environment variables appear. Before Docker 29.0 a link also copied the target's environment, passwords included, into the client as OLD_ENV_* variables; that copy is off by default in 29.x (DOCKER_KEEP_DEPRECATED_LEGACY_LINKS_ENV_VARS=1 in dockerd's environment restores it until removal in 30.0), so on older hosts treat existing links as a secret leak. Links on the default bridge are scheduled for removal in 30.0; on a user-defined network a link still works and acts as an alias. To migrate, put both containers on a user-defined network and use the container name or a --network-alias.

Rootless Docker builds its networks differently (the daemon runs in its own user and network namespace and reaches the outside through a userspace network stack); "Rootless Docker" (Advanced container security) covers what changes. Clean up:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker rm -f lab-web lab-api1 lab-api2 lab-cache lab-iv1 lab-iv2 lab-old docker network rm lab-front lab-back lab-ipvl
lab-web lab-api1 lab-api2 lab-cache lab-iv1 lab-iv2 lab-old lab-front lab-back lab-ipvl
Quick check
01A container on the user-defined network app resolves other containers by name. A teammate runs a new container with no --network flag and it cannot resolve them, although it reaches the internet. What explains both observations?
Incorrect — The resolver is wired into the namespace when the container starts; nothing waits.
Correct — Upstream resolvers explain working internet names; with no 127.0.0.11, container names are unknown to it.
Incorrect — Applications resolve names through their own libraries; a missing tool does not stop resolution.
Incorrect — Containers created later on the same network resolve names normally; this one is simply not on it.
02You run a database container on an --internal network and the application container on that network plus a normal bridge network. What can the database container do?
Incorrect — Routes are per namespace; the database has no default route of its own and does not borrow another container's.
Incorrect — Members of an internal network talk to each other; the lab connected to lab-cache on port 6379 through lab-back.
Incorrect — Neither works: outside names do not resolve on an internal network and there is no default route.
Correct — An internal network has no gateway route or masquerading, so its members reach only each other.
03docker exec app nslookup db worked until a firewall script inside app flushed the container's NAT table. Now nothing in app resolves db, although both containers are running and still share their network. Why?
Correct — Port 53 on 127.0.0.11 is redirected by NAT rules inside the container's namespace to the ports dockerd listens on.
Incorrect — Network membership is Docker's record and the veth pair; a flushed table inside the container does not change it.
Incorrect — dockerd keeps listening; the queries no longer reach its port.
Incorrect — Records live in dockerd; the NAT rules only redirect the DNS traffic to it.

Try this

Work through “Legacy links” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from networks, drivers and dns, keep “Legacy links”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related