Production best practices

A pre-ship checklist for images and the containers that run them.

Intermediate13 min · lesson 23 of 24
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (8 files, 3 KB): bestpractices.tar.gz. The lab VM shares no folders with your computer, so fetch them inside the VM: cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-hard/bestpractices.tar.gz && tar -xzf bestpractices.tar.gz, which creates ~/lab/bestpractices/. SHA-256: ad259697f48c3fafa5d876ba70cba76e912d724eb74ec68aa54a9860032f12f0

Run ./preship-check.sh against a container started with a plain docker run, and nine of fifteen checks fail: writable root filesystem, every default capability, no memory, CPU or process limit, no init process, no restart policy, logs that grow until the disk is full. The image was fine. The run was not. This lesson is a pre-ship checklist for both halves, applied to a small Node.js service, with a command behind every item. Each item has an owner lesson that explains the mechanism; here the point is to apply all of them together and check them the same way every time. It runs on the main lab VM, secopslog-docker, in ~/lab/bestpractices with the lesson files.

The image

The service is the payments API: Node's built-in http module, one small dependency (ms), a /healthz route, and a SIGTERM handler that closes the server and exits 0. The lesson files include package.json, the package-lock.json that npm ci requires, server.js, healthcheck.js and a .dockerignore that keeps .git, .env files and a local node_modules out of the build context. The Dockerfile:

Dockerfile
# node:24-alpine, pinned to the index digest the tag pointed at when this lab was written.
# Renovate or Dependabot moves the pin forward through a reviewed pull request.
FROM node:24-alpine@sha256:ebfe2f90462722a7a4de65e91990e97fe0d401c70e0e762c5b53302f905ec1c1 AS deps
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci --omit=dev
FROM node:24-alpine@sha256:ebfe2f90462722a7a4de65e91990e97fe0d401c70e0e762c5b53302f905ec1c1
# The app never runs a package manager. Remove npm, npx, corepack and yarn so they are
# not shipped, scanned and patched with every release (the base layer still holds the bytes).
RUN rm -rf /usr/local/lib/node_modules /opt/yarn-v* \
/usr/local/bin/npm /usr/local/bin/npx /usr/local/bin/corepack /usr/local/bin/yarn /usr/local/bin/yarnpkg
WORKDIR /app
ENV NODE_ENV=production PORT=3000
COPY --from=deps /app/node_modules ./node_modules
COPY package.json server.js healthcheck.js ./
# UID/GID 1000 is the image's "node" user. A numeric USER lets Kubernetes
# runAsNonRoot verify it without reading /etc/passwd.
USER 1000:1000
EXPOSE 3000
HEALTHCHECK --interval=10s --timeout=3s --start-period=5s --retries=3 \
CMD ["node", "healthcheck.js"]
CMD ["node", "server.js"]

Start with the pin. Both stages name node:24-alpine by tag and index digest. Get the digest from the registry instead of copying it from an old README:

ubuntu@secopslog-docker:~/lab/bestpractices · Docker 29.8.2
$ grep -m1 '^FROM' Dockerfile docker buildx imagetools inspect node:24-alpine --format '{{.Manifest.Digest}}'
FROM node:24-alpine@sha256:ebfe2f90462722a7a4de65e91990e97fe0d401c70e0e762c5b53302f905ec1c1 AS deps sha256:ebfe2f90462722a7a4de65e91990e97fe0d401c70e0e762c5b53302f905ec1c1

The pin matched the tag when this lab ran. When the Node maintainers rebuild 24-alpine, the second line changes and the Dockerfile keeps building from the pinned bytes until a reviewed change moves the pin; "Pinning, SBOMs, provenance and scanning" (Advanced container security) covers automating that with Renovate or Dependabot. Node 24 is a long-term-support line; in October 2026 it hands the active-LTS role to Node 26 and moves to maintenance, so check the Node release schedule when you pick or move a line. Pin the index digest, not one platform's manifest, so arm64 and amd64 builders each get their own image.

The deps stage copies only package.json and package-lock.json, then runs npm ci --omit=dev, which installs exactly the lockfile's versions and leaves development dependencies out. Because the source is copied later, editing server.js does not invalidate the install. Build once (with a fake token in .env to test the ignore file), change the source, and rebuild. The change here is a comment with the current time appended to server.js, so it is new every time you run it:

ubuntu@secopslog-docker:~/lab/bestpractices · Docker 29.8.2
$ echo 'API_TOKEN=lab-fake-token' > .env docker build -q -t lab-ship-api:1.4.2 .
sha256:00f4200c81b75c43f378b5f167040afb7012593d991238e535f19ac5926ea098
$ echo "// changed $(date -u +%FT%TZ)" >> server.js docker build --progress=plain -t lab-ship-api:1.4.2 . 2>&1 | grep -E '^#[0-9]+ (\[|CACHED)'
#1 [internal] load build definition from Dockerfile #2 [internal] load metadata for docker.io/library/node:24-alpine@sha256:ebfe2f90462722a7a4de65e91990e97fe0d401c70e0e762c5b53302f905ec1c1 #3 [internal] load .dockerignore #4 [internal] load build context #5 [deps 1/4] FROM docker.io/library/node:24-alpine@sha256:ebfe2f90462722a7a4de65e91990e97fe0d401c70e0e762c5b53302f905ec1c1 #6 [stage-1 3/5] WORKDIR /app #6 CACHED #7 [stage-1 2/5] RUN rm -rf /usr/local/lib/node_modules /opt/yarn-v* /usr/local/bin/npm /usr/local/bin/npx /usr/local/bin/corepack /usr/local/bin/yarn /usr/local/bin/yarnpkg #7 CACHED #8 [deps 2/4] WORKDIR /app #8 CACHED #9 [deps 3/4] COPY package.json package-lock.json ./ #9 CACHED #10 [deps 4/4] RUN npm ci --omit=dev #10 CACHED #11 [stage-1 4/5] COPY --from=deps /app/node_modules ./node_modules #11 CACHED #12 [stage-1 5/5] COPY package.json server.js healthcheck.js ./

Every step down to COPY --from=deps is CACHED, including RUN npm ci; only the last COPY ran again. On a real project with hundreds of packages, that ordering is the difference between a ten-second and a three-minute build. "Layers and the build cache" (Docker for beginners) explains the cache rules.

The runtime stage deletes npm, npx, corepack and yarn. The app never calls them, and the npm bundled in the base image carries its own dependency tree: when this lesson was written, Trivy 0.75.0 reported seven HIGH findings with fixes available in node_modules/npm, none in the app. Deleting them in a later layer does not shrink the image (the base layer still holds the bytes), but the files are gone from the container's filesystem, so scanners and anyone inside the container no longer find them. A distroless Node base goes further ("Minimal bases: distroless, scratch, static and Alpine" in Advanced container security).

ubuntu@secopslog-docker:~/lab/bestpractices · Docker 29.8.2
$ docker image inspect lab-ship-api:1.4.2 \ --format 'User={{.Config.User}}{{println}}Healthcheck={{json .Config.Healthcheck.Test}}{{println}}Cmd={{json .Config.Cmd}}' docker run --rm --entrypoint ls lab-ship-api:1.4.2 -a /app docker run --rm --entrypoint sh lab-ship-api:1.4.2 -c 'command -v npm || echo "no npm in the image"'
User=1000:1000 Healthcheck=["CMD","node","healthcheck.js"] Cmd=["node","server.js"] . .. healthcheck.js node_modules package.json server.js no npm in the image

Three properties to check on every image. User is 1000:1000, numeric, which is the image's node account; Kubernetes runAsNonRoot can verify a number but not a name. The health check runs node healthcheck.js, a four-line script using Node's built-in fetch, so it needs no curl or wget. That matters as soon as someone switches the base to a Debian slim variant to escape an Alpine musl issue:

ubuntu@secopslog-docker:~/lab/bestpractices · Docker 29.8.2
$ docker run --rm debian:trixie-slim sh -c 'command -v curl wget || echo "neither curl nor wget"'
neither curl nor wget

A wget health check copied from an Alpine image, where BusyBox provides it, fails on every probe there, and the container sits unhealthy while serving traffic perfectly. The listing of /app shows no .env, because the .dockerignore kept the fake token out. And the runtime image has no npm.

The run

An image cannot set its own limits or drop its own capabilities. That is the run configuration, here a Compose file:

compose.yaml
name: lab-ship
services:
api:
build: .
image: lab-ship-api:1.4.2
ports:
- "127.0.0.1:3000:3000"
read_only: true
tmpfs:
- /tmp:size=16m,mode=1777
cap_drop: [ALL]
security_opt:
- no-new-privileges:true
init: true
mem_limit: 256m
memswap_limit: 256m
cpus: 1.0
pids_limit: 100
restart: unless-stopped
logging:
driver: local
options:
max-size: 10m
max-file: "3"

read_only with a small tmpfs for /tmp; cap_drop: [ALL], since a Node server on port 3000 needs no capability; no-new-privileges so no setuid binary can raise privileges; init: true so Docker's docker-init runs as PID 1, forwards signals and reaps zombies; memory with memswap_limit equal to mem_limit, so the container cannot use swap beyond its memory limit; a CPU quota and a process limit; restart: unless-stopped so a crash or a host reboot brings it back while a deliberate stop stays stopped; and the local log driver with size-based rotation. Start it and run the checklist script from the lesson files:

ubuntu@secopslog-docker:~/lab/bestpractices · Docker 29.8.2
$ docker compose up -d --wait
... Container lab-ship-api-1 Healthy
$ ./preship-check.sh lab-ship-api-1
OK user 1000:1000 OK privileged false OK host-namespaces none OK host-mounts none OK devices none OK read-only-rootfs true OK cap-drop ALL, added back: none OK no-new-privileges set OK memory 268435456 bytes, no swap OK cpus 1000000000 nano-CPUs OK pids 100 OK init docker-init OK healthcheck healthy OK restart-policy unless-stopped OK log-rotation local 10m 0 check(s) failed

preship-check.sh reads the container's effective settings with docker inspect and exits 1 if any check fails, so a deploy job can gate on it. Besides the items in the table it fails a container that shares a host namespace (network_mode: host, pid: host and the like), bind-mounts the Docker socket or a sensitive host path, or gets host devices. It is a common subset, not a full review: capabilities added back after cap_drop: [ALL] are listed for a human to judge, seccomp and AppArmor profiles are not checked, and a log driver other than json-file, local or none passes because its rotation happens in the system it sends to. "Host namespace sharing as attack surface" and "The big three: privileged, docker.sock and host mounts" in Advanced container security explain why those settings are on the list.

The script reads configuration. The process inside the container shows what the kernel actually applied:

ubuntu@secopslog-docker:~/lab/bestpractices · Docker 29.8.2
$ docker compose exec api sh -c 'id; grep -E "^(CapEff|NoNewPrivs)" /proc/self/status; ps -o pid,user,args'
uid=1000(node) gid=1000(node) groups=1000(node) CapEff: 0000000000000000 NoNewPrivs: 1 PID USER COMMAND 1 node /sbin/docker-init -- docker-entrypoint.sh node server.js 7 node {MainThread} node server.js 28 node ps -o pid,user,args
$ docker compose exec api node -e "require('fs').writeFileSync('/tmp/cache.json', '{}'); console.log('/tmp ok')" docker compose exec api node -e "require('fs').writeFileSync('/app/cache.json', '{}')" 2>&1 | grep -m1 EROFS echo "exit=${PIPESTATUS[0]}"
/tmp ok Error: EROFS: read-only file system, open '/app/cache.json' exit=1
$ docker compose exec api sh -c "cd /sys/fs/cgroup && grep -H . memory.max memory.swap.max pids.max cpu.max"
memory.max:268435456 memory.swap.max:0 pids.max:100 cpu.max:100000 100000

CapEff is all zeros and NoNewPrivs is 1 for a process started in the container. PID 1 is docker-init, with Node as its child. A write to /tmp works, and a write under /app fails with EROFS (exit=1 is node's exit status). The cgroup files show 256 MiB of memory, no swap, 100 processes and one CPU (100000 microseconds of every 100000).

Watch out
--read-only fails at run time, not at build time, and often not at startup. An app that writes a cache file, a PID file or uploaded data to its own directory starts fine and fails on the first request that writes, with an EROFS: read-only file system error like the one above. Before shipping, exercise the paths that write, and give each of them a tmpfs (scratch data) or a volume (data that must survive). "Read-only root filesystem" (Advanced container security) shows how to find them.

The same image, run plainly

ubuntu@secopslog-docker:~/lab/bestpractices · Docker 29.8.2
$ docker run -d --name lab-plain lab-ship-api:1.4.2 >/dev/null until [ "$(docker inspect -f '{{.State.Health.Status}}' lab-plain)" = healthy ]; do sleep 1; done ./preship-check.sh lab-plain
OK user 1000:1000 OK privileged false OK host-namespaces none OK host-mounts none OK devices none FAIL read-only-rootfs false FAIL cap-drop drop='none' FAIL no-new-privileges not set FAIL memory unlimited FAIL cpus unlimited FAIL pids unlimited FAIL init no init process OK healthcheck healthy FAIL restart-policy none FAIL log-rotation json-file without max-size 9 check(s) failed

Same image, same healthy status, nine failures. The image-level items pass because they live in the image. Everything else is the default of docker run: unlimited memory, CPU and processes, the default capability set, a writable root, json-file logs without rotation and no restart. Run the check in the deploy job, not once by hand.

When a limit is hit

A memory limit turns a leak into a restart instead of a host-wide out-of-memory event that kills processes at random. From the operator's side, this is what it looks like:

ubuntu@secopslog-docker:~/lab/bestpractices · Docker 29.8.2
$ docker run -d --name lab-oom --memory 32m --memory-swap 32m --restart on-failure:3 \ alpine:3.22 sh -c 'tail /dev/zero' until [ "$(docker inspect -f '{{.State.Status}}' lab-oom)" = exited ]; do sleep 1; done docker inspect lab-oom --format 'OOMKilled={{.State.OOMKilled}} ExitCode={{.State.ExitCode}} RestartCount={{.RestartCount}}'
421cd44ef2e86eea3a020b83ffc267d12dcaf8d7b03fd20538896b04e7106df7 OOMKilled=true ExitCode=137 RestartCount=3

tail /dev/zero reads a stream with no newline and buffers it until the 32 MiB limit. OOMKilled=true and exit code 137 (128 + 9, SIGKILL) are the kernel's out-of-memory killer at work, and RestartCount=3 with the container now exited is on-failure:3 giving up. The logs of such a container usually say nothing, because a SIGKILL cannot be caught. Read the state, then raise the limit or fix the leak; "cgroups v2 and resource containment" (Advanced container security) covers how the kernel picks and what memory.high adds. Set --memory-swap equal to --memory, as here: with --memory alone, Docker allows the same amount again in swap on hosts that have swap, and the container slows down long before it is killed.

Finally, the reason for exec-form CMD, the SIGTERM handler and the init process:

ubuntu@secopslog-docker:~/lab/bestpractices · Docker 29.8.2
$ time docker compose stop api
Container lab-ship-api-1 Stopping Container lab-ship-api-1 Stopped real 0m0.233s user 0m0.055s sys 0m0.028s

A fraction of a second. A process that ignores SIGTERM would hold every stop, deploy and node drain for the full ten-second timeout before Docker sends SIGKILL.

ubuntu@secopslog-docker:~/lab/bestpractices · Docker 29.8.2
$ docker rm -f lab-plain lab-oom docker compose down docker image rm lab-ship-api:1.4.2 rm .env
... Untagged: lab-ship-api:1.4.2 Deleted: sha256:587695a45be1b77c7533d248dfd0395ef2f5e0ba9b5bdb5bc2f1aa166d105cce

What this checklist does not cover is where the image came from and whether it has known vulnerabilities. That belongs in the pipeline, before anything reaches a registry: "Docker in CI/CD: build to promote" puts the scan, SBOM, signature and promotion in order.

Quick check
01A team switches FROM node:24-alpine to a Debian slim Node base to avoid a musl problem. The service answers requests normally, but docker ps shows it unhealthy permanently. The Dockerfile still has HEALTHCHECK CMD wget -qO- http://127.0.0.1:3000/healthz || exit 1. Most likely cause?
Incorrect — Loopback works the same on any base; the probe never gets as far as connecting.
Incorrect — A slow start would delay healthy, not prevent it forever.
Incorrect — HEALTHCHECK runs independently of the CMD form.
Correct — Probe with the runtime you know is there, such as node healthcheck.js.
02preship-check.sh reports memory 268435456 bytes but memory+swap=536870912. What does that mean on a host with swap enabled?
Correct — Set the swap limit (memswap_limit / --memory-swap) equal to the memory limit to forbid swap.
Incorrect — The memory limit applies; swap is allowed in addition to it.
Incorrect — RAM use is still capped at 256 MiB; the extra allowance is swap, which slows the service before any kill.
Incorrect — Limits are ceilings, not reservations.
03A container keeps restarting and docker logs shows only normal startup lines before each restart. docker inspect shows OOMKilled=true ExitCode=137. What is the next step?
Incorrect — Exit 137 is SIGKILL from the OOM killer, which no handler can catch.
Incorrect — The restart policy only reacts to the exit; the cause is the memory limit.
Correct — The kernel killed it at its memory limit; either the app leaks or the limit is too low for its real working set.
Incorrect — Nothing was dropped; a SIGKILL gives the process no chance to write anything.

Try this

Work through “When a limit is hit” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from production best practices, keep “When a limit is hit”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related