Catch secrets before commit with gitleaks and pre-commit
Block credentials at the git hook, scan history for what already leaked, and wire the same check into CI.
git add config/prod.env && git commit -m "prod config"Detect hardcoded secrets.................................................Failed- hook id: gitleaks- exit code: 1nothing was written to history; the key is still only in the working tree. The hook ran the command in the next blockgitleaks git --pre-commit --redact --staged --verbose --exit-code 1 .Finding: AWS_ACCESS_KEY_ID=REDACTEDRuleID: aws-access-tokenFile: config/prod.envLine: 4Fingerprint: config/prod.env:aws-access-token:4INF 0 commits scanned.WRN leaks found: 1--redact prints REDACTED (a partial view is --redact=50). The same command with the key in the working tree but not staged scanned nothing and exited 0A secret that reaches a commit has to be rotated, because a clone, a fork, a CI cache or a mirror may already hold it, and rewriting history removes it from none of those. The cheap outcome is the one above: the hook ran on the staged diff, the commit was refused, and the only cleanup is deleting a line. gitleaks is the scanner in all three places this post covers, and since v8.19 its commands are git, dir and stdin; detect and protect still run but are deprecated, and most copied snippets use them.
The hook: pinned, and easy to bypass by design
repos:- repo: https://github.com/gitleaks/gitleaksrev: v8.30.1 # a tag; pre-commit autoupdate bumps it in a reviewable diffhooks:- id: gitleaks # upstream entry: gitleaks git --pre-commit --redact --staged --verbose
The upstream hook definition already passes --pre-commit --staged --redact, so the id alone is the whole configuration, and adding args duplicates or contradicts it. pre-commit install writes .git/hooks/pre-commit once per clone, which is the hook's limit: a developer who never ran it, a machine set up last year, or SKIP=gitleaks git commit (a documented, legitimate bypass for a known false positive) all skip it. The hook is for fast feedback at the keyboard. Enforcement lives in the next section, and treating the hook as enforcement is the mistake that makes teams surprised when CI finds a key. One detail for anyone testing the hook: the access key id from the AWS documentation, AKIAIOSFODNN7EXAMPLE, is allowlisted by the default rules and produces no finding, so a test with it proves nothing; use a made-up id with the same shape.
CI is the gate, with the same version and config
secret_scan:stage: testimage:name: ghcr.io/gitleaks/gitleaks:v8.30.1entrypoint: [""]variables:GIT_DEPTH: 0 # the git command needs history, not a shallow clonescript:- gitleaks git --redact --exit-code 1 --report-format sarif --report-path gitleaks.sarif--log-opts="$CI_MERGE_REQUEST_DIFF_BASE_SHA..HEAD" .artifacts:when: alwayspaths: [gitleaks.sarif]rules:- if: $CI_PIPELINE_SOURCE == "merge_request_event"
--log-opts narrows the scan to the commits the merge request adds, which keeps the job fast and keeps a decade of history from failing every pipeline. The image tag and the hook's rev should match, and a .gitleaks.toml in the repository root is picked up by both automatically, so a rule tuned for the hook behaves the same in CI. Required on protected branches, this job is what makes --no-verify pointless. GitLab's own Secret Detection template is the alternative for teams that want the findings in the security dashboard rather than as an artifact; it runs a different engine, so keep one or the other as the gate.
gitleaks git --redact --exit-code 1 --log-opts="main~1..HEAD" --report-format sarif --report-path gitleaks.sarif --verbose .Finding: AWS_ACCESS_KEY_ID=REDACTEDRuleID: aws-access-tokenFile: config/prod.envLine: 4Commit: 55cc5b0668ff7d711be3853f80426b87dc6f52f6Fingerprint: 55cc5b0668ff7d711be3853f80426b87dc6f52f6:config/prod.env:aws-access-token:4INF 2 commits scanned.WRN leaks found: 1gitleaks git --redact --exit-code 1 --log-opts="HEAD~1..HEAD" .INF 1 commits scanned.INF no leaks foundexit 1 then exit 0: the range decides. A range that starts after the offending commit finds nothing, which is why the job uses the merge request base, not HEAD~1History: scan once, baseline, then only new findings
gitleaks git --redact --report-path gitleaks-baseline.json .WRN leaks found: 23jq -r ".[] | [.RuleID, .File, .Date[:10]] | @tsv" gitleaks-baseline.json | sort | uniq -c | sort -rn | head -3 11 generic-api-key test/fixtures/keys.json 2021-03-02 7 aws-access-token scripts/deploy-old.sh 2019-11-14 5 private-key docker/dev-cert.pem 2020-06-30twenty-three findings is a typical first run on an old repository; the fixture history below holds onegitleaks git --redact --exit-code 1 --report-path gitleaks-baseline.json .; echo exit=$?WRN leaks found: 1exit=1gitleaks git --redact --exit-code 1 --baseline-path gitleaks-baseline.json .; echo exit=$?INF 3 commits scanned.INF no leaks foundexit=0git add config/ci.env && git commit -qm "ci config" # a token committed after the baselinegitleaks git --redact --exit-code 1 --baseline-path gitleaks-baseline.json --verbose .RuleID: github-patFile: config/ci.envINF 4 commits scanned.WRN leaks found: 1the baseline is a normal report; anything in it is ignored, anything new is notTwenty-three findings on a first run is normal, and they make a triage list rather than twenty-three incidents. Fixtures and test keys get a gitleaks:allow comment on the line, or a fingerprint in .gitleaksignore, each with a reason in the commit that adds it. The two waivers have different scopes, and the run showed the difference: a fingerprint taken from a git scan carries the commit hash (55cc5b06…:config/prod.env:aws-access-token:4) and silences that finding in history scans, but a gitleaks dir scan of the working tree produces the same finding with a path-only fingerprint and reports it again; the gitleaks:allow comment silences both, and --ignore-gitleaks-allow on the CI job brings it back on demand. The real keys are assumed compromised and rotated in the system that issued them (IAM, the Git host, Vault), and the ticket for each records the date the key first appeared, which the report gives you. History rewriting is a separate decision for repositories that must be published; it shrinks the exposure surface and does not undo it.
Recovering from a false positive or a lost key, in order
| Situation | Do | Do not |
|---|---|---|
| the hook blocks a test fixture | gitleaks:allow on the line with a reason in the commit; or SKIP=gitleaks git commit once, and let the CI job (same config) confirm it | git commit --no-verify as a habit: the CI job catches it and the merge request goes red anyway |
| a real key reached a commit | rotate it where it was issued first; then remove the line, and record the fingerprint in .gitleaksignore only after rotation so the scan history stays honest | rewrite history as the fix: every clone, fork and cache still has the key |
| the baseline hid a finding it should not have | git checkout <sha> -- gitleaks-baseline.json to the previous report, or delete the entry from the JSON array; the finding fails the next run | regenerate the baseline on every pipeline: it absorbs every new key |
Scanning finds secrets that should not be in Git; it does not decide where they should be. Values that belong next to the code go in encrypted with SOPS and age, and the ones CI needs at runtime come from the pipeline's secret store or an OIDC exchange, so there is nothing for the scanner to find.