BuildKit builds: stages, cache and mounts
Multi-stage builds, parallel stages, cache, bind, secret and SSH mounts.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-hard/multistage.tar.gz && tar -xzf multistage.tar.gz, which creates ~/lab/multistage/. SHA-256: 85bf2a0cb775d41db90dfc231af5bef0e3f0830e5cf3f9e94c95027dbb517167A Go service ships as a 492MB image. The binary inside it is 5MB. The rest is the Go toolchain, Alpine's package manager and the source tree, all copied in because the Dockerfile compiles and runs in the same image. Every pull moves those bytes, every scanner reports their CVEs, and every code change recompiles the whole standard library because one COPY . . invalidated the cache. This lesson fixes all three with BuildKit, the builder behind docker build: multi-stage builds, stages that run in parallel, and RUN --mount for caches, sources, secrets and SSH agents.
Use the main lab VM, secopslog-docker. The Go module, the Dockerfiles and .dockerignore are in the lesson files download; unpack them into ~/lab/multistage. No host settings change.
BuildKit is already the builder
Docker Engine 29 builds with BuildKit by default. docker build is handled by the Buildx plugin, which sends the build to a builder; on a fresh install that is the default builder, BuildKit running inside dockerd:
docker in the DRIVER column means the builder lives in the daemon and writes straight into the local image store. BuildKit v0.33.1 is the version embedded in Engine 29.8.2. "Buildx builders, outputs, cache and Bake" adds builders that run in their own containers; everything in this lesson uses the default one. The legacy builder still exists behind DOCKER_BUILDKIT=0, is deprecated, and does not understand any of the RUN --mount options below.
Stages and targets
Each FROM starts a stage. A stage can start from an image or from an earlier stage, and COPY --from=<stage> copies files out of another stage. Only the stage you build (by default the last one) becomes the image. This Dockerfile has five:
FROM golang:1.27-alpine@sha256:8a5910f31396cd4d89662f56c68b3ae31d374308270a1c3bd96672ee5ed43414 AS baseWORKDIR /srcENV CGO_ENABLED=0FROM base AS buildRUN --mount=type=bind,target=. \--mount=type=cache,id=lab-go-build,target=/root/.cache/go-build \go build -trimpath -ldflags="-s -w" -o /out/lab-app .FROM base AS testRUN --mount=type=bind,target=. \--mount=type=cache,id=lab-go-build,target=/root/.cache/go-build \go test ./... && touch /tmp/test.okFROM base AS vetRUN --mount=type=bind,target=. \--mount=type=cache,id=lab-go-build,target=/root/.cache/go-build \go vet ./... && touch /tmp/vet.okFROM scratch AS checksCOPY --from=test /tmp/test.ok /COPY --from=vet /tmp/vet.ok /FROM scratch AS runtimeCOPY --from=build /out/lab-app /lab-appUSER 65532:65532ENTRYPOINT ["/lab-app"]
base holds what every Go step needs. build, test and vet each start from base and do not depend on each other. checks collects marker files from test and vet. runtime starts from scratch, an empty filesystem, and receives one file from build. The golang image is pinned by tag and digest as in "Tags, digests and promotion", and the final stage runs as the numeric user 65532 because scratch has no /etc/passwd to look a name up in ("Run as non-root" in Advanced container security covers numeric IDs). The RUN --mount lines are explained in the next sections. Build the default target:
Read the step names. BuildKit ran [base 1/2], [base 2/2], [build 1/1] and [runtime 1/1], and nothing from test, vet or checks. It builds the graph of stages the target needs and skips the rest, so a test stage in the same Dockerfile costs nothing in a plain image build. [build 1/1] took 10.7s, almost all of it compiling the standard library packages that net/http pulls in. Your times depend on the CPU; the order of magnitude is what matters here. For comparison, the single-stage version that most projects start with:
FROM golang:1.27-alpine@sha256:8a5910f31396cd4d89662f56c68b3ae31d374308270a1c3bd96672ee5ed43414WORKDIR /srcCOPY . .RUN CGO_ENABLED=0 go build -trimpath -ldflags="-s -w" -o /usr/local/bin/lab-app .ENTRYPOINT ["lab-app"]
Same source, same binary, 7.73MB against 492MB on disk (DISK USAGE counts the unpacked layers, CONTENT SIZE the compressed ones; "Inspecting, exporting and cleaning up images" explains both columns). The builder stages still exist in the build cache on this machine, but nothing in lab-app:1 refers to them, so they are never pushed. scratch works here because the binary is static (CGO_ENABLED=0); choosing between scratch, distroless and Alpine is the subject of "Minimal bases: distroless, scratch, static and Alpine" in Advanced container security.
Cache mounts versus the layer cache
The layer cache ("Layers and the build cache" in Docker for beginners) is all or nothing per step: if a step's inputs changed, the step runs again from scratch. In Dockerfile.single, any edit changes COPY . ., so go build starts with an empty Go build cache every time. Make a one-line change and rebuild both versions:
17.6 seconds for the single-stage rebuild, 2.0 seconds for the multi-stage one. The difference is --mount=type=cache,id=lab-go-build,target=/root/.cache/go-build on the RUN line. A cache mount is a directory BuildKit keeps between builds and mounts into the step while it runs. The step still re-ran (its source changed), but Go found the compiled standard library from the previous build in its own cache and only recompiled main. The mount's contents never become part of a layer, so the image does not grow. id names the cache; steps that use the same id share it, which is why the test and vet stages below reuse the compiler output. Without id, the target path is the id.
Cache mounts are build cache records of type exec.cachemount, 90.83MB here, and docker builder prune removes them like any other build cache. Treat them as an accelerator, not storage: a cold CI runner starts with an empty one, and the build must still be correct. For package managers that lock their cache (apt, for example), add sharing=locked so parallel builds wait for each other instead of corrupting it; the default is sharing=shared.
The other mount on those lines, --mount=type=bind,target=., makes the build context visible at the working directory for the duration of the step, read-only by default. Nothing is copied into a layer, so the source tree never ends up in the build stage's filesystem, and the step's cache key still covers the files it can see.
Stages in parallel
test and vet both depend only on base. Build the checks target, which needs both. -o type=cacheonly tells BuildKit to run the build without exporting an image, since a stage of marker files is not worth keeping; output types are covered in "Buildx builders, outputs, cache and Bake". time measures the whole build:
Step #7 (vet) and step #8 (test) both start, print ... while the other one is logging, and finish within 0.2 seconds of each other at about 22.5 seconds. The whole build took 23.0 seconds, not the 45 that running them one after the other would take. BuildKit runs every step whose inputs are ready, so independent stages are free parallelism. In CI this is the usual layout: one Dockerfile, a checks (or lint, test) target the pipeline builds first, and the release target afterwards, with both sharing the base and the cache.
The build context and .dockerignore
Before any step runs, the client sends the build context (the directory given as the last argument) to the builder. "Writing a Dockerfile" in Docker for beginners introduced .dockerignore to keep secrets out of COPY . .; it also decides how much is sent. Add the two classic offenders, a Git object store and a log file, and compare a build with and without the ignore file:
Each build prints two transfers. The first (#3) is BuildKit reading .dockerignore, the file and a little metadata, and almost nothing once the file is renamed away. The second is the context. BuildKit transfers it incrementally: it sends only files that changed since the last build plus metadata for the rest, so a rebuild of an unchanged directory moves only a few dozen bytes. With the ignore file, the new .git pack and debug.log were excluded, and the context transfer stayed at that metadata-sized figure (86 bytes here). Without it, both new files went up: 35.01MB. On a laptop next to the daemon that is a fraction of a second; to a remote builder it is an upload on every build, and with --mount=type=bind,target=. the same filter decides what the build can see. The patterns follow Go's filepath.Match rules with ** support, and ! re-includes a path. A file named <Dockerfile>.dockerignore next to a Dockerfile (for example Dockerfile.single.dockerignore) overrides the shared one for builds of that Dockerfile.
Build metadata for the pipeline
CI jobs need the digest of what they just built, to scan it, sign it and deploy it. Do not parse it from the progress output. --metadata-file writes it as JSON:
containerimage.digest is the image index digest, the value the later pipeline stages should carry ("Tags, digests and promotion"). buildx.build.ref identifies the build record (builder/node/ID), which docker buildx history can look up later. The file also holds a summary of the provenance BuildKit recorded; its materials list the base images by digest. Provenance and SBOM attestations as build outputs are introduced in "Buildx builders, outputs, cache and Bake" and taught in "Pinning, SBOMs, provenance and scanning" in Advanced container security.
Secret mounts
A build that needs a token (a private package index, an API to fetch a licence file) must not receive it as a build argument or environment variable; those land in the image metadata or its history. "Build-time secrets and how images leak them" in Advanced container security shows each leak. The mechanism is a secret mount. The lab uses a dummy token in a file:
FROM alpine:3.22@sha256:5291449c3df73caf6ed85e649dec1b9e818b39a5d8c871e97afc13e9cd5e8fa8RUN --mount=type=secret,id=api_token,required=true \echo "token file: $(wc -c < /run/secrets/api_token) bytes"RUN --mount=type=secret,id=api_token,env=API_TOKEN,required=true \echo "API_TOKEN has ${#API_TOKEN} characters in this RUN"RUN ls /run/secrets 2>&1; echo "API_TOKEN=${API_TOKEN:-<unset>}"
--secret id=api_token,src=... makes the file available to the build under the id api_token. In the first RUN it is a file at /run/secrets/api_token; target= would put it elsewhere. In the second, env=API_TOKEN exposes it as an environment variable for that one command. The third RUN has no mount, and neither the file nor the variable exists there. Three behaviours are worth seeing before you rely on this in CI:
First, the rotated token (passed from an environment variable with env=LAB_TOKEN this time) did not re-run anything; every step is CACHED. Secret contents are deliberately not part of the cache key, so a build that must react to a new secret needs --no-cache or a changed input. Second, a build with no secret at all succeeded, because the steps came from cache and never asked for it; required=true is enforced only when the step actually runs. Third, with --no-cache the same build fails with secret api_token: not found, as it should. Finally, the image:
History records the commands, not the mount options or values, and /run/secrets does not exist in the image. What the step writes to its own filesystem is another matter: a RUN that copies the secret into a file, or prints it into a log the build keeps, leaks it like any other content.
SSH mounts
Cloning a private Git repository during a build needs an SSH key, and a copied key file is a leaked key. --ssh forwards an SSH agent instead: the build can ask the agent to sign, but never sees the key. The lab creates a throwaway key with no passphrase and a temporary agent; never use a real key for this exercise.
FROM alpine:3.22@sha256:5291449c3df73caf6ed85e649dec1b9e818b39a5d8c871e97afc13e9cd5e8fa8RUN apk add --no-cache openssh-clientRUN --mount=type=ssh \echo "SSH_AUTH_SOCK=$SSH_AUTH_SOCK" && ssh-add -lRUN ssh-add -l; echo "exit status without the mount: $?"
--ssh default forwards the agent named by SSH_AUTH_SOCK in your shell. Inside the step with --mount=type=ssh, SSH_AUTH_SOCK points at a socket BuildKit provides, and ssh-add -l lists the same fingerprint as the key on the VM (fingerprints differ on every run, since the key is new). The next step, without the mount, has no agent. --ssh default=./lab_ed25519 would load an unencrypted key file without an agent. For a real clone the step also needs the server's host key, otherwise ssh refuses to connect:
RUN apk add --no-cache git openssh-client# known_hosts holds GitHub's published host keys, committed to the repositoryCOPY --chmod=0644 known_hosts /root/.ssh/known_hostsRUN --mount=type=ssh git clone git@github.com:example-org/private-lib.git /src/private-lib
Many examples run ssh-keyscan github.com in the build instead. That trusts whatever answers at build time, which is exactly the check a host key exists for. Copy the keys from GitHub's published fingerprints page (docs.github.com, "GitHub's SSH key fingerprints") into a known_hosts file once, review it like any other change, and copy it in as above.
Clean up
The docker builder prune filter removes the Go cache mount by its id, so other build cache on the VM stays. /opt/secopslog-lab/setup/reset-lab.sh --images removes everything lab-related if something was left behind.
deps, build, lint and release, where release copies from build and lint starts from deps. CI runs docker build -t app:1.4 .. Which stages run?--secret id=npm_token,src=token.txt. After rotating the token you rebuild; the build is green and the RUN --mount=type=secret,id=npm_token npm ci step shows CACHED. Was the new token used, and why?Try this
Work through “Clean up” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from buildkit builds: stages, cache and mounts, keep “Clean up”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.