Multi-stage Docker builds that cut image size by 80%

Split build tooling from the runtime so your final image ships only the binary and its deps — smaller, faster, safer.

Jun 24, 2026·Updated ·7 min readIntermediate·By SecOpsLog · command-tested

The image that runs in production only needs the compiled program, and yet the single-stage Dockerfile most projects start with ships the compiler, the module cache, the source tree, the apt lists and, if anyone ever passed a token through ARG, that token's layer too. A multi-stage build separates building from shipping: the first stage has every tool and every secret, the last stage receives the one artifact you copy into it, and nothing else crosses the boundary.

bash — observed: the same Go service, two Dockerfiles (Docker Engine 28.5.2, linux/arm64; sizes are for this platform)observed
docker build -q -t app:single -f Dockerfile.single . && docker build -q -t app:multi -f Dockerfile.multi --secret id=netrc,src=netrc .
for i in app:single app:multi; do docker image inspect $i -f "{{.Size}}"; done
940603164 # app:single, 940.6 MB: golang:1.27-bookworm plus the source and the binary
8127801 # app:multi, 8.1 MB: the distroless base plus the binary
docker history --no-trunc --format "{{.Size}} {{.CreatedBy}}" app:multi | head -4
0B ENTRYPOINT ["/app"]
0B USER nonroot:nonroot
5.9MB COPY /out/app /app # buildkit
261kB bazel build //common:cacerts_debian13_arm64_tar
the runtime image has one layer of yours; the thirteen below it are the distroless base (certificates, tzdata, passwd, netbase). The compiler stage is not in it, and neither is anything it did. The multi-stage size is identical run to run; the single-stage one moves by a few hundred bytes between builds (940,603,771 in the clean-clone rerun), because the build cache it ships is not deterministic

What crosses the stage boundary, exactly

Only the paths named in COPY --from=<stage> end up in the final image. Environment variables, ARG values, files written anywhere else, and every layer of the build stage stay in the build cache and are not part of what docker push sends. That is also the trap: COPY --from=build / / copies the entire build filesystem and defeats the point. On the engine used for this article it did not even get that far: the whole-root copy onto a distroless base failed with cannot replace to directory …/var/lock with file, because the two filesystems disagree about what /var/lock is. The version of the mistake that does build is copying whole directories, and it costs what you would expect: COPY --from=build /src /src plus COPY --from=build /usr/local/go /usr/local/go produced a 246.6 MB image with main.go and the go binary in its layers. Copy the binary, or the dist/ directory, by explicit path.

Dockerfile.multi
# syntax=docker/dockerfile:1
# --- build stage: toolchain, module cache, secrets allowed ---
FROM golang:1.27 AS build
WORKDIR /src
COPY go.mod go.sum ./
RUN --mount=type=cache,target=/go/pkg/mod go mod download
COPY . .
RUN --mount=type=cache,target=/go/pkg/mod \
--mount=type=secret,id=netrc,target=/root/.netrc \
CGO_ENABLED=0 go build -trimpath -ldflags="-s -w" -o /out/app ./cmd/app
# --- test stage: runs in CI, never becomes the shipped image ---
FROM build AS test
RUN go vet ./... && go test ./...
# --- runtime stage: the binary, CA certificates, tzdata, nothing else ---
FROM gcr.io/distroless/static-debian13:nonroot
COPY --from=build /out/app /app
USER nonroot:nonroot
ENTRYPOINT ["/app"]

Three details in that file do real work. RUN --mount=type=secret exposes a credential to one build step without writing it into any layer, which ARG and ENV cannot promise. The test stage inherits from build and is selected with docker build --target test in CI, so tests run against the exact compiled artifact while the default build (no --target) produces the runtime image without test binaries. And CGO_ENABLED=0 is what makes the static distroless variant usable: a dynamically linked binary needs base-debian13 (glibc and libssl) or cc-debian13 instead.

In CI the two invocations share a cache: docker build --target test --secret id=netrc,src=$HOME/.netrc . compiles and tests, then docker build -t app:$SHA . reuses every layer of the build stage from the first command and only adds the runtime stage. The secret is passed at build time from the runner's file and is readable only inside the one RUN that mounts it; docker history app:$SHA afterwards shows no layer that could contain it, which is the check worth automating. A stronger version of that check is to search the layers themselves rather than their descriptions, and it is the one the fixture runs: docker save the image, extract every layer tar, and grep the raw content for the token.

bash — observed: the secret is in no layer, and the search would have found itobserved
layer_grep() { d=$(mktemp -d); docker save "$1" | tar -x -C "$d"; n=0; for b in "$d"/blobs/sha256/*; do n=$((n + $(tar -xOf "$b" 2>/dev/null | grep -c -- "$2"))); done; echo "$n"; }
layer_grep app:multi p2c-fake-netrc-token-never-issued
0
printf "FROM gcr.io/distroless/static-debian13:nonroot\nCOPY netrc /netrc\n" | docker build -q -t control -f - . && layer_grep control p2c-fake-netrc-token-never-issued
1
the control image copies the same fake netrc into a layer and the same search finds it, so the 0 above means absent, not unsearched. The same listing shows no cmd/app/main.go and no /usr/local/go path in the runtime image

Which stages actually run depends on the builder. BuildKit builds only the stages the selected target depends on, so a Dockerfile can carry a test stage, a lint stage and a docs stage without any of them costing a production build a second; the legacy builder processed every stage up to the target regardless. The fixture carries a never stage whose only instruction is exit 1: the default build succeeded and its progress log names only the build and runtime stages, --target test ran go vet and go test, and --target never failed the build as it should. That difference is also why a stage nobody references makes a poor hiding place: one builder skips it, the other builds and caches it where docker history can read it. A stage that must never ship is kept out of the runtime stage's COPY --from list, which is the only boundary that holds in both.

When a build change ships the wrong thing: recovery in order of preference

SituationDoDo not
a runtime image is found to contain a credential, the source tree or a toolchaintreat the credential as exposed and rotate it first; fix the COPY --from paths; rebuild and re-run the layer search before pushing; redeploy from the previous known-good digest (image: …@sha256:<last-good>) until thenpush a new tag over it and move on: the registry, every pull cache and every node that pulled it still hold the layer
the runtime image will not start after a stage change (the binary needs libc, a certificate bundle, a timezone)roll back to the previous digest, then move the runtime stage to the base that has what the binary needs (base-debian13 for a dynamically linked binary) or fix the build flags (CGO_ENABLED=0 for static)add a shell to debug it in place: debug tags belong in non-production environments
a --secret build works on a laptop and fails in CIthe secret is a file on the runner, so check that the job writes it before the build and passes the same id=; docker build --progress=plain shows which RUN mounts itfall back to ARG TOKEN=…: it lands in the layer history
What was run for this article
Docker Engine 28.5.2 with BuildKit on linux/arm64, golang:1.27-bookworm (go1.27.1) and gcr.io/distroless/static-debian13:nonroot pulled by digest, against a 40-line Go HTTP service with one test. The terminal blocks marked observed are copied from that run: the two sizes, the history, the layer search with its positive control, the per-directory copy at 246.6 MB, the whole-root copy that failed to build, and the three --target outcomes. Fourteen exit codes are asserted. The sizes are for this architecture and these base images and will differ elsewhere. The CI cache sharing across two jobs, the registry push and the Node and Java stages are representative and were not executed; the recovery table follows the documented model.
Verify the image, not the Dockerfile
docker history and an image scanner show what actually shipped. A COPY of a whole directory, a test fixture that contains a credential, or a base image that quietly gained a shell in an update are all visible there and invisible in the Dockerfile you remember writing.

Choosing the runtime stage

Runtime bases for a multi-stage build

BaseContainsUse when
gcr.io/distroless/static-debian13CA certificates, tzdata, /etc/passwd, no libcstatically linked Go or Rust binaries
gcr.io/distroless/base-debian13the above plus glibc and libsslcgo, or any dynamically linked binary
gcr.io/distroless/nodejs24-debian13Node runtime, no npm or shella built dist/ plus production node_modules
nginx:alpine or a static servera shell and a package managerserving compiled front-end assets; accept the larger surface
scratchnothinga static binary that needs no certificates or timezone data

For interpreted stacks the artifact is a directory rather than a binary. A node:24 build stage runs npm ci && npm run build; the runtime stage copies dist/ and a production-only node_modules produced by npm ci --omit=dev in a separate step, not the development tree from the build stage. Java is the same shape with a JDK stage feeding a java21-debian13 runtime.

The size drop is the visible result; the security result is that the tools an attacker relies on after code execution (a shell, curl, a package manager) are not present. What distroless removes and what it leaves is the subject of running containers without a shell, and the layer ordering that keeps these builds fast is in Dockerfile layer caching.

Go deeper in a courseAdvanced container securityMinimal images, non-root runtime, seccomp and the escape paths a small image closes.View course

Related posts

Quick reference