Scanning container images with Trivy in GitLab CI
Wire Trivy into your pipeline, fail builds on criticals without blocking every merge, and cache the vuln DB so scans stay under 30 seconds.
Whether an image contains known vulnerabilities is not in doubt; every base image does. The question a pipeline decides is where you learn about them: in the merge request that introduced the base image, or in the incident channel after it has been the default deploy target for a month. The difference is entirely in job ordering. A scan that runs after docker push is a report; a scan that the push job depends on is a control.
Two scan passes read the same image. The report pass exits 0 and lists everything; the gate pass exits 1 on CRITICAL only. Push depends on the gate through needs:, so a failed gate leaves the registry untouched.
trivy image --input image.tar --severity LOW,MEDIUM,HIGH,CRITICAL --exit-code 0 -q | grep ^Total; echo "exit ${PIPESTATUS[0]}"Total: 49 (LOW: 14, MEDIUM: 26, HIGH: 6, CRITICAL: 3)exit 0trivy image --input image.tar --severity CRITICAL --ignore-unfixed --exit-code 1; echo "exit $?"INFO Detected OS family="alpine" version="3.18.0"INFO [alpine] Detecting vulnerabilities... os_version="3.18" repository="3.18" pkg_num=15WARN This OS version is no longer supported by the distributionimage.tar (alpine 3.18.0)Total: 3 (CRITICAL: 3)│ busybox │ CVE-2022-48174 │ CRITICAL │ fixed │ 1.36.0-r9 │ 1.36.1-r1 │ busybox: stack overflow vulnerability in ash.c leads to ││ busybox-binsh │ │ │ │ │ │ arbitrary code execution ││ ssl_client │ │ │ │ │ │ │exit 1same image, same database: the report pass lists 49 and exits 0, the gate pass lists the three fixable CRITICALs (one CVE in three busybox packages) and exits 1, so the push job never runs for this commit. At HIGH the gate would also be red (six fixable)The pipeline
Build saves the image as a tar artifact so the scan jobs read exactly the bytes that will be pushed, without pulling from a registry they are supposed to be protecting. The build job is shown with docker; on runners without a Docker socket, buildah build followed by buildah push … oci:image/ produces an OCI layout that trivy image --input reads the same way. The scanner image is pinned to a version; a floating tag on the tool that enforces policy is a supply-chain hole of its own. The vulnerability database is pulled as an OCI artifact from mirror.gcr.io/aquasec first and ghcr.io/aquasecurity second; caching TRIVY_CACHE_DIR under a project-wide key means one download per DB update instead of one per job, which also keeps a busy runner fleet away from registry rate limits.
variables:IMAGE: "$CI_REGISTRY_IMAGE:$CI_COMMIT_SHORT_SHA"TRIVY_CACHE_DIR: .trivycache/stages: [build, scan, push]build:stage: buildscript:- docker build -t "$IMAGE" .- docker save "$IMAGE" -o image.tarartifacts:paths: [image.tar]expire_in: 1 day.trivy:stage: scanimage:name: aquasec/trivy:0.74.0entrypoint: [""]cache:key: trivy-dbpaths: [.trivycache/]needs: [build]scan-report:extends: .trivyscript:- trivy image --input image.tar --severity LOW,MEDIUM,HIGH,CRITICAL --exit-code 0- trivy image --input image.tar --format cyclonedx --output sbom.cdx.jsonartifacts:paths: [sbom.cdx.json]scan-gate:extends: .trivyscript:- trivy image --input image.tar --severity CRITICAL --ignore-unfixed --exit-code 1push:stage: pushneeds: [build, scan-gate]script:- docker load -i image.tar- docker push "$IMAGE"
Why CRITICAL only, and why two passes
A gate that blocks on every HIGH in a freshly adopted base image is disabled within a month, usually with allow_failure: true and a comment promising to revisit. The report pass keeps the full list visible in every pipeline so the backlog is known; the gate pass blocks only on what a developer can fix now, so --ignore-unfixed is on it. Tightening is a per-project decision to make when the backlog is short: move HIGH into the gate pass for a service once its report has been clean at that level for a few weeks, not for the whole organisation on a date.
Two things make the gate a control rather than a report. scan-gate carries no allow_failure, so a red gate fails the pipeline and the push stage never starts. And push names the gate in needs, which makes the dependency explicit rather than a side effect of stage order: a job with needs runs only when every job it lists has succeeded, however the stages are rearranged later. When the gate is red, the image tagged with that commit exists only as an artifact that expires in a day.
cat .trivyignore.yamlvulnerabilities: - id: CVE-2022-48174 statement: accepted until the base image bump, INFRA-442 expired_at: 2027-01-01trivy image --input image.tar --severity CRITICAL --ignore-unfixed --exit-code 1 --ignorefile .trivyignore.yaml; echo "exit $?"INFO Some vulnerabilities have been ignored/suppressed. Use the "--show-suppressed" flag to display them.│ image.tar (alpine 3.18.0) │ alpine │ 0 │ - │exit 0sed -i "s/2027-01-01/2024-01-01/" .trivyignore.yaml && trivy image --input image.tar --severity CRITICAL --ignore-unfixed --exit-code 1 --ignorefile .trivyignore.yaml -q | grep ^Total; echo "exit ${PIPESTATUS[0]}"Total: 3 (CRITICAL: 3)exit 1the exception is a file in the repository, reviewed like code, and it lapses on its own; a bad YAML value (a colon inside an unquoted statement) made Trivy exit 1 with an error, which the gate treats the same as a findingWhen the gate is wrong: recovery in order of preference
| Situation | Do | Do not |
|---|---|---|
| a CRITICAL with a fix that the base image does not carry yet | .trivyignore.yaml entry with expired_at a few weeks out and the ticket in statement, in the same merge request | allow_failure: true on scan-gate: needs still passes and the push runs on every red gate from then on |
| the gate is red on a hotfix that must ship | push from a pipeline that runs the report pass and records the exception first; the hotfix carries the .trivyignore.yaml change, and the next scheduled registry scan re-opens it | retag the previous image by hand: the registry now holds an image no pipeline scanned |
| the database download fails and every job is red | set TRIVY_SKIP_DB_UPDATE=true on the gate job with the cached .trivycache/ (the last good database) until the mirror recovers; keep the report pass running so the outage is visible | --exit-code 0 on the gate to get green: the control is gone and nothing says so |
The same two flags drive the dependency scan on the source tree one stage earlier, where the suppression file with reasons and expiry dates is described. After the gate, the natural next control is to sign the image that passed, so that admission in the cluster can distinguish an image this pipeline vouched for from one that merely has the same name.
Go deeper in a courseSecure CI/CD with GitLabScanning, signing, SBOMs and policy gates as one pipeline.View course