Recipe: NGINX as a reverse proxy

TLS edge, static files, proxying, and re-resolving names after a backend moves.

Intermediate14 min · lesson 21 of 24
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (7 files, 2 KB): r-nginx.tar.gz. The lab VM shares no folders with your computer, so fetch them inside the VM: cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-hard/r-nginx.tar.gz && tar -xzf r-nginx.tar.gz, which creates ~/lab/r-nginx/. SHA-256: 9716bc4f668be731d61fb17d86a620220f31a94dfc423597aa715d0e8d5ef7d1

A backend container restarts during a deploy, comes back healthy, and every request through the proxy now answers 502 Bad Gateway until somebody reloads NGINX. The NGINX error log has the reason: connect() failed (111: Connection refused) while connecting to upstream, upstream: "http://172.18.0.2:8000/". That address no longer belongs to the backend. This recipe builds the usual edge for a Compose stack (NGINX terminating TLS, serving static files and proxying /api/ to a private backend), reproduces that 502, and fixes it with Docker's DNS server. It runs on the main lab VM, secopslog-docker, in ~/lab/r-nginx with the lesson files.

The stack

Two services. web is nginxinc/nginx-unprivileged, the NGINX image built to run as UID 101 and to listen on 8080 and 8443. api is a twenty-line Python server that echoes what the proxy sent it, so you can see the headers arrive:

compose.yaml
name: lab-edge
services:
web:
image: nginxinc/nginx-unprivileged:1.30-alpine@sha256:15c994d10d6d78658721c3bcafff14cb281fba2a4bdf9d5ba92c416a472516e3
ports:
- "127.0.0.1:80:8080"
- "127.0.0.1:443:8443"
volumes:
- ./${EDGE_CONF:-default.conf}:/etc/nginx/conf.d/default.conf:ro
- ./site:/usr/share/nginx/html:ro
secrets: [tls_cert, tls_key]
read_only: true
tmpfs:
- /tmp
cap_drop: [ALL]
security_opt: [no-new-privileges:true]
healthcheck:
test: ["CMD", "wget", "-q", "-O", "/dev/null", "http://127.0.0.1:8080/healthz"]
interval: 10s
timeout: 3s
retries: 3
start_period: 5s
depends_on:
api:
condition: service_healthy
restart: unless-stopped
api:
image: python:3.14-slim@sha256:f85c5697265c178cc6887276c55fe16cf3d14ca35c3df6a5eab3b360534a55d2
command: ["python", "/app/api.py"]
init: true
user: "65534:65534"
volumes:
- ./api/api.py:/app/api.py:ro
read_only: true
cap_drop: [ALL]
security_opt: [no-new-privileges:true]
healthcheck:
test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/healthz', timeout=2)"]
interval: 10s
timeout: 3s
retries: 3
start_period: 5s
restart: unless-stopped
secrets:
tls_cert:
file: ./certs/fullchain.pem
tls_key:
file: ./certs/privkey.pem

Only web publishes ports, and only on the VM's loopback address; api has no ports: at all, so the one way to it is through the proxy on the Compose network. Both images are pinned by tag and digest. Both containers run with a read-only root filesystem, every capability dropped and no-new-privileges. NGINX only needs /tmp writable: the unprivileged image keeps its PID file and temporary request bodies there, hence the tmpfs. init: true on api puts a small init process in front of Python so docker compose stop takes a fraction of a second instead of the ten-second kill timeout ("CMD, ENTRYPOINT and PID 1" (Docker for beginners) explains why). The certificate and key arrive as Compose secrets. Both services have health checks; web checks only that NGINX answers its own /healthz, so a dead backend does not mark the proxy unhealthy. That is a choice: if you would rather have the orchestrator replace the proxy when the backend is gone, probe through it instead.

The NGINX configuration:

default.conf
# Docker's embedded DNS server. Re-check answers every 10 s instead of
# trusting the record's 600 s TTL.
resolver 127.0.0.11 valid=10s ipv6=off;
resolver_timeout 3s;
server {
listen 8080;
location = /healthz { access_log off; return 200 "ok\n"; }
location / { return 301 https://$host$request_uri; }
}
server {
listen 8443 ssl;
ssl_certificate /run/secrets/tls_cert;
ssl_certificate_key /run/secrets/tls_key;
location / {
root /usr/share/nginx/html;
try_files $uri $uri/ /index.html;
}
location /api/ {
# A variable in proxy_pass makes NGINX resolve "api" at request time.
set $api_upstream http://api:8000;
# With a variable, proxy_pass sends the URI unchanged, so strip /api here.
rewrite ^/api/(.*)$ /$1 break;
proxy_pass $api_upstream;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
}

The plain-HTTP server answers /healthz and redirects everything else to HTTPS. The TLS server serves files from /usr/share/nginx/html and forwards /api/ to the backend with four headers: the original Host, the client address in X-Real-IP and X-Forwarded-For, and X-Forwarded-Proto: https, so the application knows the client used TLS even though the hop inside Docker is plain HTTP. The resolver line, the variable and the rewrite are the fix for the 502; the section after next explains them. ssl_protocols is not set, because the default since NGINX 1.23.4 is already TLS 1.2 and 1.3.

The key nobody can read

Create a self-signed certificate for localhost (a real deployment uses one from your CA or Let's Encrypt; the configuration is the same):

ubuntu@secopslog-docker:~/lab/r-nginx · Docker 29.8.2
$ mkdir certs openssl req -x509 -newkey ec -pkeyopt ec_paramgen_curve:P-256 -nodes -days 30 \ -keyout certs/privkey.pem -out certs/fullchain.pem \ -subj "/CN=localhost" -addext "subjectAltName=DNS:localhost" 2>/dev/null ls -l certs
total 8 -rw-r--r-- 1 ubuntu ubuntu 607 Oct 8 01:11 fullchain.pem -rw------- 1 ubuntu ubuntu 241 Oct 8 01:11 privkey.pem

OpenSSL writes the private key as 0600, owned by ubuntu. Start the stack:

ubuntu@secopslog-docker:~/lab/r-nginx · Docker 29.8.2
$ docker compose up -d --wait
... container lab-edge-web-1 is unhealthy
$ docker compose logs --no-log-prefix web | grep -m1 emerg
2026/10/07 19:41:56 [emerg] 1#1: cannot load certificate key "/run/secrets/tls_key": BIO_new_file() failed (SSL: error:8000000D:system library::Permission denied:calling fopen(/run/secrets/tls_key, r) error:10080002:BIO routines::system lib)

--wait waits for the health checks and exits 1 because web never became healthy. Outside Swarm mode, Compose hands a secret to the container as a read-only bind mount of the host file, so it keeps the host owner (UID 1000) and mode (0600). NGINX runs as UID 101, which cannot open it, and exits at startup. The fix is to make the file readable inside the container and protect it on the host with the directory instead:

ubuntu@secopslog-docker:~/lab/r-nginx · Docker 29.8.2
$ chmod 700 certs chmod 644 certs/privkey.pem ls -ld certs certs/privkey.pem
drwx------ 2 ubuntu ubuntu 4096 Oct 8 01:11 certs -rw-r--r-- 1 ubuntu ubuntu 241 Oct 8 01:11 certs/privkey.pem
$ docker compose up -d --wait
... Container lab-edge-web-1 Healthy
$ docker compose ps --format 'table {{.Name}}\t{{.Status}}\t{{.Ports}}'
NAME STATUS PORTS lab-edge-api-1 Up 14 seconds (healthy) lab-edge-web-1 Up 5 seconds (healthy) 127.0.0.1:80->8080/tcp, 127.0.0.1:443->8443/tcp

Mode 0644 on the key looks wrong until you look at the directory: 0700 means no other host user can reach the file, while inside the container the bind mount exposes the file alone. Inside the container, though, 0644 lets every process read the key, whatever its user. If you can change the file's group, chgrp 101 (the image's nginx group) with mode 0640 narrows that to the NGINX user; the lab keeps 0644 because it is the setting that works without root on the host. The long-syntax uid, gid and mode fields on a Compose secret look like the cleaner fix, and they are ignored outside Swarm; "Recipe: WordPress and MySQL with secrets" shows the warning Compose prints. The ps output confirms the topology: api has no published port.

Checking each path

ubuntu@secopslog-docker:~/lab/r-nginx · Docker 29.8.2
$ curl -sI http://localhost/ | grep -E '^(HTTP|Location)'
HTTP/1.1 301 Moved Permanently Location: https://localhost/
$ curl -s --cacert certs/fullchain.pem https://localhost/
<!doctype html> <html lang="en"> <head><title>lab-edge</title></head> <body><p>Static page served by NGINX.</p></body> </html>
$ curl -s --cacert certs/fullchain.pem "https://localhost/api/orders?page=2"
{ "served_by": "172.18.0.2", "path": "/orders?page=2", "host": "localhost", "x_forwarded_for": "172.18.0.1", "x_forwarded_proto": "https" }

Plain HTTP gets a 301 to HTTPS. The static page comes off disk over TLS, and --cacert makes curl verify the self-signed certificate instead of skipping verification with -k. The API reply is the backend describing the request it received: the path arrived as /orders?page=2, without the /api prefix and with the query string intact. host is still localhost, and x_forwarded_proto is https. x_forwarded_for is 172.18.0.1, the network's gateway address: curl connected to a port published on 127.0.0.1, which Docker serves through docker-proxy, the userland proxy, and the proxy opens its own connection to NGINX from the gateway. A remote client reaching a port published on all addresses is DNATed and keeps its own address, unless the userland proxy handles it too ("Publishing ports and the packet path" covers both paths). Addresses on the Compose network vary between machines and runs.

One more claim worth checking, because it is repeated widely: that only root can bind ports below 1024, so non-root NGINX must use 8080:

ubuntu@secopslog-docker:~/lab/r-nginx · Docker 29.8.2
$ docker compose exec web sh -c 'id; sysctl net.ipv4.ip_unprivileged_port_start'
uid=101(nginx) gid=101(nginx) groups=101(nginx) net.ipv4.ip_unprivileged_port_start = 0

Rootful Docker sets net.ipv4.ip_unprivileged_port_start=0 in every container's network namespace, so UID 101 could bind port 80 here. The unprivileged image listens on 8080 by convention, and that convention also works under rootless Docker and in Kubernetes with restricted policies. Run NGINX as non-root because it is non-root, not because of the port. "Capabilities, cap-drop and no-new-privileges" (Advanced container security) covers the sysctl.

Why the proxy keeps a dead address

Most NGINX examples write proxy_pass http://api:8000/;. static-upstream.conf is that version. The Compose file mounts whichever file EDGE_CONF names, so start web with it:

ubuntu@secopslog-docker:~/lab/r-nginx · Docker 29.8.2
$ EDGE_CONF=static-upstream.conf docker compose up -d --wait
... Container lab-edge-web-1 Healthy
$ curl -s --cacert certs/fullchain.pem https://localhost/api/ | grep served_by
"served_by": "172.18.0.2",

Now give the backend a new address. Recreating a container usually gets the same address back, because the old one is released first. It changes when something else takes the address while the backend is down: a one-off job, another service scaling up, a rolling deploy. Simulate that with a job container:

ubuntu@secopslog-docker:~/lab/r-nginx · Docker 29.8.2
$ docker compose stop api docker run -d --name lab-edge-job --network lab-edge_default alpine:3.22 sleep 600 docker compose start --wait api
Container lab-edge-api-1 Stopping Container lab-edge-api-1 Stopped 740b80c9d3c8633844e5e3cb68ad99c5f1b553e25e8c165a09bf717153928ea8 Container lab-edge-api-1 Starting Container lab-edge-api-1 Started Container lab-edge-api-1 Waiting Container lab-edge-api-1 Healthy
$ docker inspect -f '{{.Name}} {{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' lab-edge-job lab-edge-api-1
/lab-edge-job 172.18.0.2 /lab-edge-api-1 172.18.0.4
$ curl -s -o /dev/null -w '%{http_code}\n' --cacert certs/fullchain.pem https://localhost/api/ docker compose logs --no-log-prefix web | grep -m1 'connect() failed'
502 2026/10/07 19:42:19 [error] 21#21: *4 connect() failed (111: Connection refused) while connecting to upstream, client: 172.18.0.1, server: , request: "GET /api/ HTTP/1.1", upstream: "http://172.18.0.2:8000/", host: "localhost"

api is healthy at 172.18.0.4, and the proxy answers 502. With a literal hostname, NGINX resolves api once, when it loads the configuration, and keeps that address until the next reload. It is still sending requests to 172.18.0.2, which now belongs to the job container. Here the job does not listen on 8000, so the connection is refused. If it did, NGINX would send your API traffic, with its headers and cookies, to whatever container holds the old address. Docker's DNS record for api was correct the whole time.

The fix in default.conf has three parts. resolver 127.0.0.11 points NGINX at Docker's embedded DNS server, which every container on a user-defined network can reach at that address ("Networks, drivers and DNS"). valid=10s caps how long NGINX trusts an answer; Docker's DNS answers with a TTL of 600 seconds, which would otherwise apply. And proxy_pass takes a variable: NGINX resolves names in a variable at request time through resolver, instead of once at startup. The variable changes one more behaviour. With a variable, proxy_pass passes the request URI through unchanged, and the trailing-slash prefix stripping of the literal form no longer happens, so the rewrite ... break line removes /api itself. A side benefit: NGINX starts even when api does not exist yet, where the literal form fails at startup with host not found in upstream.

Put the fixed configuration back and move the backend again:

ubuntu@secopslog-docker:~/lab/r-nginx · Docker 29.8.2
$ docker compose up -d --wait
... Container lab-edge-web-1 Healthy
$ curl -s --cacert certs/fullchain.pem https://localhost/api/ | grep served_by
"served_by": "172.18.0.4",
$ docker compose stop api docker run -d --name lab-edge-job2 --network lab-edge_default alpine:3.22 sleep 600 docker compose start --wait api
Container lab-edge-api-1 Stopping Container lab-edge-api-1 Stopped 012a067cfbf78ac1f9f075d1dac03bdc07f44a3c8d2477829ece4547db52ebff Container lab-edge-api-1 Starting Container lab-edge-api-1 Started Container lab-edge-api-1 Waiting Container lab-edge-api-1 Healthy
$ for i in 1 2 3 4 5 6 7; do curl -s -o /dev/null -w '%{http_code} ' --cacert certs/fullchain.pem https://localhost/api/ sleep 2 done; echo curl -s --cacert certs/fullchain.pem https://localhost/api/ | grep served_by docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' lab-edge-api-1
200 200 200 200 200 200 200 "served_by": "172.18.0.5", 172.18.0.5

Every request succeeded and reached the backend at its new address. The backend's restart took longer than ten seconds here (the health check runs every ten seconds), so NGINX's cached answer had expired by the first request. A move faster than valid can still get 502 or 504 responses for up to ten seconds; lower valid if that matters, at the cost of more DNS lookups. NGINX 1.27.3 and later also offer server api:8000 resolve; inside an upstream block with a zone, which re-resolves in the background and keeps upstream keepalive connections; use it when you need load balancing over several backend addresses.

Renewing the certificate

NGINX reads certificates when it loads its configuration. Renew the files, test the configuration, reload, and check which certificate a client now gets:

ubuntu@secopslog-docker:~/lab/r-nginx · Docker 29.8.2
$ openssl s_client -connect localhost:443 -servername localhost </dev/null 2>/dev/null | openssl x509 -noout -serial
serial=11622E7E3C3069B1B53FF184352EE7DC142836AD
$ openssl req -x509 -newkey ec -pkeyopt ec_paramgen_curve:P-256 -nodes -days 30 \ -keyout certs/privkey.pem -out certs/fullchain.pem \ -subj "/CN=localhost" -addext "subjectAltName=DNS:localhost" 2>/dev/null chmod 644 certs/privkey.pem docker compose exec web nginx -t docker compose exec web nginx -s reload
nginx: the configuration file /etc/nginx/nginx.conf syntax is ok nginx: configuration file /etc/nginx/nginx.conf test is successful 2026/10/07 19:42:50 [notice] 49#49: signal process started
$ sleep 1 openssl s_client -connect localhost:443 -servername localhost </dev/null 2>/dev/null | openssl x509 -noout -serial
serial=5342D0509FC48FC2F8882E29A341EA42D7FD8F77

Different serial numbers: the reload picked up the new certificate without dropping connections. nginx -t first, always; a reload with a broken configuration is refused, but a restart with one takes the edge down. One trap with single-file bind mounts, which Compose secrets are: OpenSSL rewrote the files in place, so the container saw the new content. A tool that writes a new file and renames it over the old one (many editors, certbot's symlink swap, sed -i) leaves the container looking at the old inode until the container is recreated. Mount a directory instead when your renewal tool works that way, and make it readable by UID 101.

Baking the configuration into an image

For production you may prefer an image that carries its configuration and site, so the image is the whole artifact. Comments in a Dockerfile must start the line; a comment after an instruction is passed to it as arguments:

ubuntu@secopslog-docker:~/lab/r-nginx · Docker 29.8.2
$ docker build -q -f Dockerfile.broken -t lab-edge:broken .
Dockerfile.broken:1 -------------------- 1 | >>> FROM nginxinc/nginx-unprivileged:1.30-alpine # runs as UID 101 2 | COPY default.conf /etc/nginx/conf.d/default.conf 3 | -------------------- ERROR: failed to build: failed to solve: dockerfile parse error on line 1: FROM requires either one or three arguments
Dockerfile
# Bake the config and the site into the image for production.
# The unprivileged variant runs as UID 101 and listens on 8080/8443.
FROM nginxinc/nginx-unprivileged:1.30-alpine@sha256:15c994d10d6d78658721c3bcafff14cb281fba2a4bdf9d5ba92c416a472516e3
COPY default.conf /etc/nginx/conf.d/default.conf
COPY site/ /usr/share/nginx/html/
ubuntu@secopslog-docker:~/lab/r-nginx · Docker 29.8.2
$ docker build -q -t lab-edge:1 . docker image inspect lab-edge:1 --format 'user={{.Config.User}}'
sha256:b636c7bffdf04d6cc86a0d70305f12f293aa06c33d84a63d4a6c39ac38c1c4e9 user=101

The certificate and key still stay out of the image: anyone who can pull an image can read every layer of it. Keep them as secrets or mounts. Clean up:

ubuntu@secopslog-docker:~/lab/r-nginx · Docker 29.8.2
$ docker rm -f lab-edge-job lab-edge-job2 docker compose down docker image rm lab-edge:1 rm -rf certs
... Untagged: lab-edge:1 Deleted: sha256:b636c7bffdf04d6cc86a0d70305f12f293aa06c33d84a63d4a6c39ac38c1c4e9
Quick check
01After a deploy, docker compose ps shows api healthy, but NGINX returns 502 for /api/ and its log shows connect() failed to an address that docker inspect says belongs to a different container. The config uses proxy_pass http://api:8000/;. What is happening?
Incorrect — Docker updated the record as soon as the new container joined the network; the stale address is cached inside NGINX.
Correct — A literal hostname is resolved at load time. Use resolver 127.0.0.11 with a variable in proxy_pass, or server ... resolve in an upstream block.
Incorrect — NGINX reaches it over the Compose network by name; publishing would only expose it to the host.
Incorrect — A failing health check does not restart a container under plain Compose, and this one probes NGINX itself, which is fine.
02You switch to set $api http://api:8000; proxy_pass $api/; and now every request for /api/orders?page=2 reaches the backend as /. Why?
Incorrect — DNS resolution has nothing to do with the request URI.
Incorrect — The cache time affects when the name is looked up again, not the URI that is sent.
Incorrect — A variable can hold a full URL with a port; the lab used exactly that.
Correct — Prefix stripping only works with a literal URI. Keep proxy_pass $api; without a path and strip /api with rewrite ^/api/(.*)$ /$1 break;.
03NGINX in nginxinc/nginx-unprivileged exits with cannot load certificate key "/run/secrets/tls_key" ... Permission denied. The key file on the host is -rw------- ubuntu ubuntu. The Compose secret already sets uid: "101" and mode: 0400. What fixes it?
Correct — Outside Swarm the secret is a bind mount of the host file with its own owner and mode, and Compose warns that uid, gid and mode are ignored.
Incorrect — Capabilities in the bounding set do not help a non-root process by default, and handing out DAC_OVERRIDE to read one file is far broader than needed.
Incorrect — Root would read it, but it gives up the reason for the unprivileged image; making the file readable is the targeted fix.
Incorrect — Then everyone who can pull the image has the private key, in a layer that never goes away.

Try this

Work through “Baking the configuration into an image” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from recipe: nginx as a reverse proxy, keep “Baking the configuration into an image”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related