Scanning Terraform with Checkov before you apply
Catch public buckets and open security groups in the plan — in CI, with suppressions that expire and a clean baseline.
The Terraform changes that cause incidents are small diffs: a bucket policy with Principal: "*", a security group rule for 0.0.0.0/0 on port 22, an RDS instance created without storage_encrypted. Checkov reads the .tf files and matches them against a few thousand built-in checks, needs no cloud credentials to do so, and exits non-zero when it finds one. The tool is easy; the decisions are which findings should stop a merge, what to do about the four hundred findings in the stack that already exists, and how an exception is recorded so it does not become permanent.
checkov -d terraform --framework terraform --compact --quietPassed checks: 14, Failed checks: 11, Skipped checks: 0Check: CKV_AWS_23: "Ensure every security group and rule has a description" FAILED for resource: aws_security_group.webCheck: CKV_AWS_23: "Ensure every security group and rule has a description" FAILED for resource: aws_security_group_rule.httpsCheck: CKV_AWS_20: "S3 Bucket has an ACL defined which allows public READ access." FAILED for resource: aws_s3_bucket.assetsplus eight more on the bucket (versioning, access logging, KMS, replication, lifecycle, public access block, event notifications, an unattached security group). One of the eleven should block the merge; the rest are worth fixing and not worth a blocked deploy. That is the design problemWhat blocks, what warns
Gate design
| Finding class | Examples | Handling |
|---|---|---|
| exposure | public bucket ACL or policy, 0.0.0.0/0 on admin ports, public RDS or Redshift | hard fail: the merge does not happen |
| data protection | unencrypted EBS, RDS or S3, logging disabled on a load balancer, no KMS on a queue that holds PII | hard fail on new resources; baseline on existing ones with a ticket |
| hygiene | missing descriptions, tags, versioning on a scratch bucket | soft fail: reported, not blocking |
| legacy debt | the 400 findings in the stack that predates the scanner | baseline file; only new findings fail |
Checkov expresses this with --hard-fail-on and --soft-fail-on, which accept check ids, and with a baseline. The two flags are not symmetrical. On the fixture, --hard-fail-on CKV_AWS_53 (a check the fixture passes) returned exit 0 with eleven failed checks still printed: once a hard-fail list exists, every check not on it is soft. --soft-fail-on CKV_AWS_20 alone still exited 1 on the other ten. So the hard-fail list is the whole gate, and a --soft-fail-on next to it is documentation for the reader rather than configuration. The flags also accept severity names, but the severity of a check comes from Prisma Cloud platform metadata and is only present when Checkov runs with an API key; on the open-source policies alone, gate by check id and by baseline rather than by a severity name that resolves to nothing.
checkov:stage: testimage: bridgecrew/checkov:3.3.17 # pinned: the policy gate is itself a dependencyscript:- checkov -d terraform/ --framework terraform--baseline terraform/.checkov.baseline # findings already known do not fail (the file lives where --create-baseline put it)--hard-fail-on CKV_AWS_20,CKV_AWS_53,CKV_AWS_54,CKV_AWS_55,CKV_AWS_56,CKV_AWS_24,CKV_AWS_25# everything not listed above is reported and does not fail the job--output cli --output gitlab_sast --output-file-path console,gl-sast-report.jsonartifacts:reports:sast: gl-sast-report.json # findings appear on the merge request diffrules:- if: $CI_PIPELINE_SOURCE == "merge_request_event"
The baseline: only new findings fail
A scanner introduced into an existing estate fails every merge on day one and is disabled on day two. checkov --create-baseline writes the current findings to a .checkov.baseline file inside the directory it scanned, not in the working directory, and that run is still a failing scan (exit 1), so it cannot sit in front of &&. From then on, --baseline compares each run against that file and fails only on findings that are not in it. The baseline is the debt register: it is committed, it is reviewed, and it should shrink. A pipeline that regenerates the baseline on every run has removed the gate while keeping the job green, which is worth a review comment when someone proposes it.
checkov -d terraform --framework terraform --create-baseline --compact --quiet; echo exit=$?Passed checks: 14, Failed checks: 11, Skipped checks: 0exit=1jq '[.failed_checks[].findings[].check_ids[]] | length' terraform/.checkov.baseline11git add terraform/.checkov.baseline && git commit -m "checkov: baseline 11 existing findings (INFRA-442 tracks the burn-down)"checkov -d terraform --framework terraform --baseline terraform/.checkov.baseline --compact --quiet; echo exit=$?exit=0nothing printed: with --quiet the eleven known findings are neither listed nor counted as failuresprintf 'resource "aws_s3_bucket" "new" {\n bucket = "acme-new-fixture"\n}\nresource "aws_s3_bucket_acl" "new" {\n bucket = aws_s3_bucket.new.id\n acl = "public-read"\n}\n' > terraform/new.tfcheckov -d terraform --framework terraform --baseline terraform/.checkov.baseline --compact --quiet; echo exit=$?Passed checks: 0, Failed checks: 8, Skipped checks: 0Check: CKV_AWS_20: "S3 Bucket has an ACL defined which allows public READ access." FAILED for resource: aws_s3_bucket.newexit=1eight findings, all on the new bucket. Note the address: the ACL finding is attributed to aws_s3_bucket.new, not to the aws_s3_bucket_acl resource that carries the acl argument, so a skip has to go on the bucketExceptions with an owner and a date
An inline checkov:skip comment suppresses one check on one resource and carries a free-text reason. It has to be inside the resource block: placed on the line above resource, where most people put comments, Checkov 3.3.17 still reports the check as FAILED; on the first line inside the braces it is honoured and the reason is echoed in the output. The reason is the part that matters: a ticket id, who accepted the risk, and when it should be looked at again. Checkov does not expire skips by itself, so the expiry is enforced by the review of the comment and by a periodic grep for skips older than their date; a skip with no reason, or a --skip-check in the pipeline definition that hides a check from every resource, is the finding to raise in review.
# CDN origin bucket: public read is intended, the origin is fronted by CloudFront with OAC and WAF.# Risk accepted by security (INFRA-442), review by 2026-12-31.resource "aws_s3_bucket" "cdn_origin" {# checkov:skip=CKV_AWS_20:public-read origin behind CloudFront OAC and WAF; INFRA-442; review 2026-12-31bucket = "acme-cdn-origin"}
checkov -d terraform --framework terraform --check CKV_AWS_20 --compact --quiet # comment above the resource blockPassed checks: 0, Failed checks: 1, Skipped checks: 0Check: CKV_AWS_20: "S3 Bucket has an ACL defined which allows public READ access." FAILED for resource: aws_s3_bucket.assetscheckov -d terraform --framework terraform --check CKV_AWS_20 --compact --quiet # comment moved inside the blockPassed checks: 0, Failed checks: 0, Skipped checks: 1Check: CKV_AWS_20: "S3 Bucket has an ACL defined which allows public READ access." SKIPPED for resource: aws_s3_bucket.assets Suppress comment: public-read origin behind CloudFront OAC and WAF; INFRA-442; review 2026-12-31exit 1, then exit 0. The echoed reason is what a reviewer greps forWhen the gate is wrong: recovery in order of preference
| Situation | Do | Do not |
|---|---|---|
| a true exception on one resource | a checkov:skip inside the block with ticket, owner and review date, in the same merge request | --skip-check in the job, which hides the check from every resource |
| a check starts failing after a Checkov upgrade and the finding is wrong | pin the job back to the last tag that judged it correctly (image: bridgecrew/checkov:<last-good>), re-run the merge request to confirm the finding is gone, open an issue upstream, move forward when it is fixed | add the check to --soft-fail-on; with a hard-fail list in place that changes nothing, and without one it hides a real finding on the next resource |
| the baseline was regenerated and swallowed a new finding | git checkout <last-good-sha> -- terraform/.checkov.baseline and re-run; the finding is reported again | delete the baseline: the four hundred legacy findings come back and the gate is disabled by lunchtime |
Rules that are specific to your organisation (a mandatory cost-center tag, an approved list of AMIs, a KMS key that production databases must use) are custom policies in YAML or Python loaded with --external-checks-dir, and they run in the same job. When the rule depends on a computed value, a variable or a module output that is only known after terraform plan, a static scan of the HCL cannot evaluate it; that case belongs to a policy on the plan JSON, which OPA on Terraform plans covers.
The scan runs before terraform plan posts its output to the merge request, so a public bucket is a red job rather than a comment someone has to read. The apply side has its own guard: remote state with locking so the apply that follows a green scan is the only one running.