Multi-stage Docker builds that cut image size by 80%
Split build tooling from the runtime so your final image ships only the binary and its deps — smaller, faster, safer.
The image that runs in production only needs the compiled program, and yet the single-stage Dockerfile most projects start with ships the compiler, the module cache, the source tree, the apt lists and, if anyone ever passed a token through ARG, that token's layer too. A multi-stage build separates building from shipping: the first stage has every tool and every secret, the last stage receives the one artifact you copy into it, and nothing else crosses the boundary.
docker build -q -t app:single -f Dockerfile.single . && docker build -q -t app:multi -f Dockerfile.multi --secret id=netrc,src=netrc .for i in app:single app:multi; do docker image inspect $i -f "{{.Size}}"; done940603164 # app:single, 940.6 MB: golang:1.27-bookworm plus the source and the binary8127801 # app:multi, 8.1 MB: the distroless base plus the binarydocker history --no-trunc --format "{{.Size}} {{.CreatedBy}}" app:multi | head -40B ENTRYPOINT ["/app"]0B USER nonroot:nonroot5.9MB COPY /out/app /app # buildkit261kB bazel build //common:cacerts_debian13_arm64_tarthe runtime image has one layer of yours; the thirteen below it are the distroless base (certificates, tzdata, passwd, netbase). The compiler stage is not in it, and neither is anything it did. The multi-stage size is identical run to run; the single-stage one moves by a few hundred bytes between builds (940,603,771 in the clean-clone rerun), because the build cache it ships is not deterministicWhat crosses the stage boundary, exactly
Only the paths named in COPY --from=<stage> end up in the final image. Environment variables, ARG values, files written anywhere else, and every layer of the build stage stay in the build cache and are not part of what docker push sends. That is also the trap: COPY --from=build / / copies the entire build filesystem and defeats the point. On the engine used for this article it did not even get that far: the whole-root copy onto a distroless base failed with cannot replace to directory …/var/lock with file, because the two filesystems disagree about what /var/lock is. The version of the mistake that does build is copying whole directories, and it costs what you would expect: COPY --from=build /src /src plus COPY --from=build /usr/local/go /usr/local/go produced a 246.6 MB image with main.go and the go binary in its layers. Copy the binary, or the dist/ directory, by explicit path.
# syntax=docker/dockerfile:1# --- build stage: toolchain, module cache, secrets allowed ---FROM golang:1.27 AS buildWORKDIR /srcCOPY go.mod go.sum ./RUN --mount=type=cache,target=/go/pkg/mod go mod downloadCOPY . .RUN --mount=type=cache,target=/go/pkg/mod \--mount=type=secret,id=netrc,target=/root/.netrc \CGO_ENABLED=0 go build -trimpath -ldflags="-s -w" -o /out/app ./cmd/app# --- test stage: runs in CI, never becomes the shipped image ---FROM build AS testRUN go vet ./... && go test ./...# --- runtime stage: the binary, CA certificates, tzdata, nothing else ---FROM gcr.io/distroless/static-debian13:nonrootCOPY --from=build /out/app /appUSER nonroot:nonrootENTRYPOINT ["/app"]
Three details in that file do real work. RUN --mount=type=secret exposes a credential to one build step without writing it into any layer, which ARG and ENV cannot promise. The test stage inherits from build and is selected with docker build --target test in CI, so tests run against the exact compiled artifact while the default build (no --target) produces the runtime image without test binaries. And CGO_ENABLED=0 is what makes the static distroless variant usable: a dynamically linked binary needs base-debian13 (glibc and libssl) or cc-debian13 instead.
In CI the two invocations share a cache: docker build --target test --secret id=netrc,src=$HOME/.netrc . compiles and tests, then docker build -t app:$SHA . reuses every layer of the build stage from the first command and only adds the runtime stage. The secret is passed at build time from the runner's file and is readable only inside the one RUN that mounts it; docker history app:$SHA afterwards shows no layer that could contain it, which is the check worth automating. A stronger version of that check is to search the layers themselves rather than their descriptions, and it is the one the fixture runs: docker save the image, extract every layer tar, and grep the raw content for the token.
layer_grep() { d=$(mktemp -d); docker save "$1" | tar -x -C "$d"; n=0; for b in "$d"/blobs/sha256/*; do n=$((n + $(tar -xOf "$b" 2>/dev/null | grep -c -- "$2"))); done; echo "$n"; }layer_grep app:multi p2c-fake-netrc-token-never-issued0printf "FROM gcr.io/distroless/static-debian13:nonroot\nCOPY netrc /netrc\n" | docker build -q -t control -f - . && layer_grep control p2c-fake-netrc-token-never-issued1the control image copies the same fake netrc into a layer and the same search finds it, so the 0 above means absent, not unsearched. The same listing shows no cmd/app/main.go and no /usr/local/go path in the runtime imageWhich stages actually run depends on the builder. BuildKit builds only the stages the selected target depends on, so a Dockerfile can carry a test stage, a lint stage and a docs stage without any of them costing a production build a second; the legacy builder processed every stage up to the target regardless. The fixture carries a never stage whose only instruction is exit 1: the default build succeeded and its progress log names only the build and runtime stages, --target test ran go vet and go test, and --target never failed the build as it should. That difference is also why a stage nobody references makes a poor hiding place: one builder skips it, the other builds and caches it where docker history can read it. A stage that must never ship is kept out of the runtime stage's COPY --from list, which is the only boundary that holds in both.
When a build change ships the wrong thing: recovery in order of preference
| Situation | Do | Do not |
|---|---|---|
| a runtime image is found to contain a credential, the source tree or a toolchain | treat the credential as exposed and rotate it first; fix the COPY --from paths; rebuild and re-run the layer search before pushing; redeploy from the previous known-good digest (image: …@sha256:<last-good>) until then | push a new tag over it and move on: the registry, every pull cache and every node that pulled it still hold the layer |
| the runtime image will not start after a stage change (the binary needs libc, a certificate bundle, a timezone) | roll back to the previous digest, then move the runtime stage to the base that has what the binary needs (base-debian13 for a dynamically linked binary) or fix the build flags (CGO_ENABLED=0 for static) | add a shell to debug it in place: debug tags belong in non-production environments |
a --secret build works on a laptop and fails in CI | the secret is a file on the runner, so check that the job writes it before the build and passes the same id=; docker build --progress=plain shows which RUN mounts it | fall back to ARG TOKEN=…: it lands in the layer history |
Choosing the runtime stage
Runtime bases for a multi-stage build
| Base | Contains | Use when |
|---|---|---|
gcr.io/distroless/static-debian13 | CA certificates, tzdata, /etc/passwd, no libc | statically linked Go or Rust binaries |
gcr.io/distroless/base-debian13 | the above plus glibc and libssl | cgo, or any dynamically linked binary |
gcr.io/distroless/nodejs24-debian13 | Node runtime, no npm or shell | a built dist/ plus production node_modules |
nginx:alpine or a static server | a shell and a package manager | serving compiled front-end assets; accept the larger surface |
scratch | nothing | a static binary that needs no certificates or timezone data |
For interpreted stacks the artifact is a directory rather than a binary. A node:24 build stage runs npm ci && npm run build; the runtime stage copies dist/ and a production-only node_modules produced by npm ci --omit=dev in a separate step, not the development tree from the build stage. Java is the same shape with a JDK stage feeding a java21-debian13 runtime.
The size drop is the visible result; the security result is that the tools an attacker relies on after code execution (a shell, curl, a package manager) are not present. What distroless removes and what it leaves is the subject of running containers without a shell, and the layer ordering that keeps these builds fast is in Dockerfile layer caching.