CI/CD interview questions
Prep CI/CD interviews from pipeline basics to secure delivery: artifacts, environments, runners, progressive deploys, OIDC, and supply-chain gates.
Levels run Beginner → Intermediate → Advanced → Expert. Answers are phrased the way you would say them in an interview; Advanced and Expert answers add the deeper reasoning, a diagram where it helps, and the follow-up an interviewer often asks next.
Fundamentals
The short version: CI merges and automatically builds and tests every change. Continuous Delivery keeps artifacts always releasable with a manual prod gate; Continuous Deployment removes that gate and ships when checks pass.
CI stops at a tested artifact; delivery adds a gate; deployment ships automatically.
Link to this questionI'd sketch build/package, test — unit, integration, lint, security scans — publish an immutable artifact, then deploy or promote through environments. Each stage should fail fast with logs you can actually read.
stages: [build, test, deploy] build: stage: build script: [make build] test: stage: test script: [make test]Link to this question
It's the immutable output of a build — jar, image, binary, package. Build once, store it with a version or digest, and promote those same bits across environments instead of rebuilding per env.
docker build -t registry/app:$CI_COMMIT_SHA . docker push registry/app:$CI_COMMIT_SHA # deploy that exact digest/tag to dev → stg → prodLink to this question
It's the worker that executes pipeline jobs — shared or dedicated, ephemeral or long-lived. I prefer ephemeral runners: clean environment per job and a smaller blast radius if a build gets compromised.
Long-lived runners are faster but they accumulate caches, credentials, and drift. Short-lived VMs or containers that tear down after the job are the default I'd argue for unless you've got a clear reason otherwise.
Link to this questionUsually a push or merge request — also tags, schedules, manual/API triggers, or an upstream pipeline. I scope jobs with rules so deploy only runs where it should.
deploy:
script: [./deploy.sh]
rules:
- if: '$CI_COMMIT_BRANCH == "main"'Link to this questionrules. The GitLab docs mark only and except as deprecated; rules checks if, changes, and exists conditions in order and can set when, variables, and allow_failure per match. I don't mix both styles in one pipeline because their defaults differ.
# deprecated
deploy_legacy:
script: [./deploy.sh]
only: [main]
# current
deploy:
script: [./deploy.sh]
rules:
- if: $CI_COMMIT_BRANCH == "main"Interviewer often follows with: What happens to a job when none of its rules match?
Link to this questionCache reuses inputs like dependencies across jobs and is best-effort. Artifacts are job outputs passed to later stages and have to be present. I'd never use cache to hand the build product to deploy.
build:
script: [make build]
artifacts:
paths: [dist/]
expire_in: 1 weekLink to this questionProtected branches — main, release — require reviews and status checks; protected environments add approvals and restrict which jobs and credentials can deploy to prod. Together they stop random branches from shipping.
deploy-prod:
stage: deploy
when: manual
environment:
name: production
rules:
- if: '$CI_COMMIT_BRANCH == "main"'Link to this questionCompile or package a single immutable artifact, then promote that same digest through dev, staging, and prod. Rebuilding per environment reintroduces drift and invalidates the earlier test evidence. Only configuration differs between environments: inject it at deploy time, and gate production with approvals and automated checks.
DIGEST=sha256:abc123 # built and tested once helm upgrade app ./chart -f values-stg.yaml --set image.digest=$DIGEST # after the production gate, the same digest: helm upgrade app ./chart -f values-prod.yaml --set image.digest=$DIGEST # value name depends on the chartLink to this question
Fail-fast stops on the first broken check to save minutes; fail-safe may run more jobs to gather signal. I'd fail-fast for expensive deploys, and sometimes continue on non-blocking linters so one MR surfaces all the issues.
trivy-fs: script: [trivy fs --exit-code 0 .] allow_failure: trueLink to this question
Pipelines in practice
I'd measure stage timings first, then cache dependencies, parallelize independent jobs, split slow suites, use path filters in monorepos, and keep images small. Optimize the longest stage, not the loudest complaint.
test:
parallel: 4
cache:
key: $CI_COMMIT_REF_SLUG
paths: [node_modules/]
script: [npm test]Link to this questionI'd detect changed paths and run only affected packages — path rules, Nx, Bazel, Turborepo — share caches, and fan out per-service jobs. Otherwise every commit rebuilds everything and people start skipping checks.
Without change detection, CI becomes the bottleneck. Use rules:changes or a build graph to select targets; keep a nightly full build for coverage gaps. Shared templates should still pin tool versions. A change in a shared lib rebuilding all consumers is correct, not a bug. I'd rather fail closed on undetected graph edges than silently skip tests.
api-test:
script: [make -C services/api test]
rules:
- changes: [services/api/**/*]Interviewer often follows with: What do you run when package.json at the repo root changes?
Link to this questionFactor shared jobs into templates (GitLab CI/CD components with spec:inputs or include/extends, GitHub reusable workflows or composite actions) so a fix lands everywhere and pipelines stay DRY.
.base-test: image: python:3.12 before_script: [pip install -r requirements.txt] unit: extends: .base-test script: [pytest -q]Link to this question
Store them in the platform secret store or Vault, inject as masked variables at runtime, scope to jobs and protected branches, never echo them, and prefer short-lived OIDC tokens over static cloud keys.
deploy:
script:
- export TOKEN=$(vault kv get -field=token secret/deploy)
# never: echo "$TOKEN"
environment: productionLink to this questionBind them from the credentials store with credentials() in environment, or a withCredentials block, and reference them inside single-quoted sh steps so the shell expands them instead of Groovy. Log masking only catches accidental prints.
pipeline {
agent any
environment {
REGISTRY = credentials('registry-creds') // also sets REGISTRY_USR and REGISTRY_PSW
}
stages {
stage('Push') {
steps {
sh 'echo "$REGISTRY_PSW" | docker login -u "$REGISTRY_USR" --password-stdin registry.example.com'
}
}
}
}Interviewer often follows with: What goes wrong if the same step uses a double-quoted Groovy string with ${REGISTRY_PSW}?
Link to this questionOne job definition runs across combinations — OS, language version, browser. It catches compatibility bugs early but multiplies minutes, so I'd keep the matrix focused on what we actually support.
strategy:
matrix:
node: [22, 24]
steps:
- uses: actions/setup-node@v7
with: { node-version: ${{ matrix.node }} }Link to this questionI'd add a render-and-validate step — kubeconform, conftest, dry-run — against the exact overlays that ship, plus a smoke test in staging that exercises the new key. Unit tests alone don't prove Kubernetes config.
CI often builds the image and skips manifest validation, or uses a different values file than prod. Fix: helm template or kustomize build of the prod overlay in CI, schema validate, policy checks, and deploy to staging with the identical chart version. Optional: kubectl apply --dry-run=server against a non-prod cluster. At runtime, fail fast on missing required env with a startup check. This is the classic pipeline-green, prod-red scenario.
helm template web charts/web -f values/prod.yaml | kubeconform -strict - helm template web charts/web -f values/prod.yaml | conftest test -
Interviewer often follows with: Would you block merge on missing keys or only block deploy?
Link to this questionI'd introduce path-filtered pipelines, remote build caches, test impact analysis, and tiered gates — always-on secret and policy scans for touched paths, full suites on release paths — so feedback is minutes, not optional.
Slow CI creates shadow IT and force-merge culture. Techniques: change detection, remote cache or container layer cache keyed safely, parallel shards, and split merge-gate vs nightly deep. Security: never skip gitleaks or OPA on path filters that could still introduce secrets or workflows; workflow files always trigger full CI security jobs. Measure p50 PR feedback time and force-merge rate. Balance speed engineering with non-negotiable security jobs.
# security job: on paths [**/.github/workflows/**, **/Dockerfile, **/*] # unit job: only when packages/foo/** changes
Interviewer often follows with: What's a safe cache keying scheme that avoids cache poisoning across forks?
Link to this questionDeployment strategies
Blue-green flips all traffic between two stacks for instant rollback at roughly double cost. Canary shifts a small percentage and watches metrics before ramping — safest when telemetry is good. Rolling replaces in place — cheapest, slower to unwind.
I pick by risk tolerance, infra budget, and observability. Stateless web apps fit all three; sticky sessions and one-shot migrations constrain canary and rolling. Always define abort criteria before you shift traffic.
I'd run blue and green Deployments behind one Service. Cutover is changing the Service selector — or Ingress weight — to green after smoke tests. Rollback points the selector back at blue; no rebuild required.
Both colors need full capacity before the flip if you want instant rollback. Keep blue scaled until green proves healthy on error rate, latency, and business KPIs. Database migrations must be backward compatible so blue can still serve after rollback; plan breaking schema changes as expand/contract migrations, separately from the flip. On Gateway API the flip can be a weight change on the HTTPRoute backendRefs instead of a selector patch. Cost is roughly 2× compute during the window. Alternatives: mesh traffic split or DNS weighted records. Document who can flip, and automate the selector patch in the pipeline with a manual approval gate for prod.
Traffic flips by selector; blue stays hot for instant rollback.
# after green is Ready and smoke-tested:
kubectl patch svc web -p '{"spec":{"selector":{"version":"green"}}}'
# rollback:
kubectl patch svc web -p '{"spec":{"selector":{"version":"blue"}}}'Interviewer often follows with: How do you handle a breaking schema change with blue-green?
Link to this questionRedeploy the previous known-good artifact or flip traffic back — blue-green or canary. Keep deploys immutable and versioned, and make DB migrations expand/contract so rollback doesn't require restoring the database.
kubectl rollout undo deploy/web helm rollback app 42 # canary: set weight back to 0% on the new versionLink to this question
They separate deploy from release: ship code dark, enable per cohort at runtime, kill-switch without redeploying. The cost is flag lifecycle — stale flags become debt.
Flags let you test in production safely, but they need ownership, defaults for failure modes, and removal after rollout. I'd never use flags as a substitute for fixing a broken pipeline.
Link to this questionUse expand/contract: additive migration first, deploy code that uses both old and new, backfill, then drop old columns in a later release. Run the migration once, as a dedicated job before traffic shifts rather than on every pod's startup, and never couple a destructive migration to a rolling deploy.
During a rolling update, old and new pods coexist. A drop-column migration breaks old pods still reading that column. Expand/contract keeps both versions compatible. Run migrations as an explicit pipeline job or PreSync hook before traffic shifts, with locks, timeouts, and a dry-run in staging, and gate the production promote on its exit code. Don't migrate on pod startup: when several replicas or canary pods boot at once, each races the same DDL. Exactly-once comes from the migration tool's lock (Flyway, Liquibase) or a job holding a lease. For high risk, separate the migrate window from the app release. Backup and PITR stay mandatory.
migrate:
script: [./migrate.sh up]
rules: [{ if: '$CI_COMMIT_BRANCH == "main"' }]
deploy:
needs: [migrate]
script: [./deploy.sh]Interviewer often follows with: What's a dual-write period and when do you end it?
Link to this questionProtected environment with required reviewers, required green checks, deploy only from main or release tags, and often a manual promotion step. Prod credentials exist only on that environment.
deploy-prod:
when: manual
environment: production
rules:
- if: '$CI_COMMIT_BRANCH == "main"'Link to this questionI'd abort the canary on the business KPI, flip traffic back, and treat golden signals as necessary but not sufficient — wire product metrics into the analysis before ramping further.
Infra metrics miss domain failures — wrong price, auth edge case, empty search. Progressive delivery needs SLIs that reflect user outcomes, not only p99. Define abort thresholds and ownership before the release. After rollback, bisect with feature flags if the artifact must stay deployed dark. Log the decision in the incident timeline so the next release reuses the same gates.
# Flagger / Rollouts style: set canary weight to 0 kubectl argo rollouts abort web kubectl argo rollouts undo web
Interviewer often follows with: Which three metrics would you require before a 50% ramp?
Link to this questionI'd look for sticky sessions, DNS TTLs, CDN caches, or multiple ingress paths that weren't flipped together — healthy green doesn't equal 100% traffic shift.
Blue-green needs an atomic traffic switch at every edge: load balancer listener, ingress, API gateway, and CDN behaviors. Long DNS TTLs and client connection pools delay cutover. Sticky cookies keep sessions on blue. Mitigations: low TTL or ALB weighted target groups, drain connections, invalidate CDN, verify with synthetics from multiple resolvers, and a hard kill switch to flip weights to 0/100. Pair with dual-write or session externalization so color affinity isn't required. Separate "deployed" from "receiving traffic."
# ALB weighted forward config → 0% blue / 100% green # curl -sI https://app.example | grep -i x-color # dig +ttl app.example # watch TTL during cutover
Interviewer often follows with: How do you drain WebSocket clients during a blue-green flip?
Link to this questionPrevious digests remain pullable, traffic shifting is automated — not a rebuild — config and DB are rollback-compatible, and the rollback path is tested. Hoping that main still builds isn't a plan.
Rollback isn't re-running the pipeline from an old commit if the registry GC'd the tag or migrations are irreversible. Design: immutable digests retained, one-click traffic revert, feature flags for dark code, and expand/contract DB discipline. Run game days that execute rollback under a timebox. Observability abort rules should trigger the same path. Artifact retention plus traffic control plus schema compatibility is the triad.
kubectl argo rollouts undo web # or: set ALB weight previous-TG=100 # confirm: digest still in registry (retention policy)
Interviewer often follows with: How do you handle rollback when the bad release already ran a destructive migration?
Link to this questionSecrets, identity & runners
The pipeline presents a short-lived, workload-bound identity token; the cloud exchanges it for temporary credentials scoped to repo and branch. Nothing long-lived to leak or rotate.
Static keys in GitHub or GitLab secrets are shared, hard to attribute, and survive forever if forgotten. OIDC ties STS issuance to claims like sub, aud, and ref. Condition the trust policy on the exact workflow file for critical roles by customizing the sub claim to include job_workflow_ref. Combine with least-privilege IAM per environment, and audit CloudTrail for AssumeRoleWithWebIdentity. Locally developers use SSO, not the CI role.
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: arn:aws:iam::123:role/ci-deploy
aws-region: us-east-1Interviewer often follows with: What claim would you use to pin a role to one workflow file?
Link to this questionPrefer ephemeral, isolated executors; no privileged Docker-in-Docker by default; pinned images; secrets not written to disk longer than needed; network egress controlled; and separate runners for untrusted vs trusted jobs.
A shared privileged runner is a lateral-movement prize — steal cloud tokens, mine crypto, or poison caches. Mitigations: Kubernetes executor or VM-per-job, drop capabilities, read-only root where possible, use runner authentication tokens instead of legacy registration tokens and rotate them, and watch for unexpected processes. Never mount the host Docker socket into untrusted jobs. Cache poisoning is real — scope caches by branch and prefer content-addressed dependencies.
# config.toml — do NOT set privileged=true for untrusted projects
[[runners]]
executor = "docker"
[runners.docker]
image = "golang:1.27"
privileged = false
disable_entrypoint_overwrite = trueInterviewer often follows with: Why is mounting /var/run/docker.sock dangerous on a runner?
Link to this questionUntrusted input — PR titles, branch names, fork code — interpolated into a shell runs with pipeline privileges. Don't run privileged jobs on untrusted PRs; pass input via env vars and quote it; isolate fork builds without secrets.
echo ${{ github.event.issue.title }} is a classic RCE vector. Bind to an environment variable, then use "$TITLE". Prefer curated actions pinned by SHA. Treat YAML from contributors as untrusted code.
- run: echo "title is $TITLE"
env:
TITLE: ${{ github.event.pull_request.title }}Interviewer often follows with: Why pin third-party actions to a full commit SHA?
Link to this questionPin to a full commit SHA (GitHub's allowed-actions policy can enforce this), review permissions requested, minimize the set allowed org-wide, and treat them as code running with your pipeline's privileges.
# mutable tag — can move uses: actions/checkout@v7 # prefer: uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1Link to this question
Set the org or repo default to read-only, declare permissions: contents: read at the top of each workflow, then grant extra scopes per job, such as id-token: write only on the job that requests an OIDC token. Once any permission is listed, every unlisted one drops to none.
permissions:
contents: read
jobs:
deploy:
runs-on: ubuntu-latest
permissions:
contents: read
id-token: write # required to request the OIDC token
steps:
- run: ./deploy.shInterviewer often follows with: Why does a pull_request_target run from a fork get a read/write token, and what does that mean for checkout?
Link to this questionMask variables in the CI platform, never echo secrets, avoid writing them into build artifacts or Docker layers, and scan logs and artifacts for high-entropy tokens in CI.
# GitLab: masked + protected variable DEPLOY_TOKEN script: - curl -H "Authorization: Bearer $DEPLOY_TOKEN" https://api/... # never: echo $DEPLOY_TOKEN | tee token.txtLink to this question
Ephemeral runners are created per job and destroyed after — clean and safer. Persistent runners reuse a machine for speed or special hardware like GPUs. Default to ephemeral; isolate and harden any persistent fleet.
Persistent runners need patching, credential scrubbing between jobs, and strong tenancy. Document why each non-ephemeral runner exists and who owns it.
Link to this questionI'd tighten the trust policy conditions on sub, ref, and workflow, remove wildcards, require environment protection, and prove it with a failing PR job that cannot AssumeRole into prod.
Common footgun: StringLike sub repo:org/name:* without ref conditions. Fix: allow only refs/heads/main and release tags, optionally pin the workflow file, and separate roles per env. GitHub Environments add required reviewers, but a job that references an environment gets sub = repo:acme/app:environment:prod, so the trust policy must match that form rather than the ref form. Verification: open a feature-branch workflow that attempts sts:AssumeRole and expect AccessDenied — keep that as a conformance test. Audit CloudTrail for unexpected assumers. Show concrete IAM condition keys, not just "we use OIDC."
"Condition": {
"StringEquals": {
"token.actions.githubusercontent.com:sub": "repo:acme/app:ref:refs/heads/main"
}
}Interviewer often follows with: How do you allow hotfix tags without allowing every tag in the repo?
Link to this questionSupply chain & artifacts
SAST, SCA or dependency scan, secret scanning, container and IaC scan, SBOM generation, and artifact signing with verification — with a failure policy that blocks criticals without drowning the team in noise.
security:
stage: test
script:
- gitleaks dir .
- trivy image --exit-code 1 registry/app:$SHA
- syft scan registry/app@$DIGEST -o cyclonedx-json=sbom.json
- cosign sign --yes registry/app@$DIGEST # DIGEST from the push stepLink to this questionBoth are machine-readable SBOM standards: SPDX is ISO/IEC 5962:2021 from the Linux Foundation, and CycloneDX is ECMA-424 from OWASP and also covers VEX. I produce what my scanners and customers consume, often both, attach it to the image digest as a signed attestation, and rescan it when new CVEs land.
syft scan registry/app@$DIGEST -o spdx-json=sbom.spdx.json -o cyclonedx-json=sbom.cdx.json cosign attest --yes --type cyclonedx --predicate sbom.cdx.json registry/app@$DIGEST grype sbom:./sbom.cdx.json
Interviewer often follows with: Why attest the SBOM to the digest instead of only storing it as a build artifact?
Link to this questionCI signs with cosign — key or keyless — after build; admission or the deploy job verifies the signature and optionally provenance against a trusted identity before apply. Unsigned or wrong-identity images fail closed.
Signing without verification is theater. Verify in two places when you can: deploy pipeline with cosign verify, and cluster admission with Kyverno or Gatekeeper plus sigstore. Keyless binds to the OIDC identity of the builder — pin that identity in policy. Prefer digests over tags. Store SBOM and attestations alongside the image. Rotate trust roots carefully with a dual-allow window.
Build signs; cluster or pipeline verifies before the workload runs.
cosign verify \ --certificate-identity-regexp 'https://github.com/org/app/.github/workflows/.*' \ --certificate-oidc-issuer https://token.actions.githubusercontent.com \ registry/app@sha256:abc... kubectl set image deploy/web app=registry/app@sha256:abc...
Interviewer often follows with: Tag mutable vs digest — which do you put in production manifests?
Link to this questionProvenance is only as trustworthy as the builder that signs it. At SLSA Build L1 the pipeline merely generates it, which is trivial to forge. L2 needs a hosted build platform that generates and signs the provenance itself; L3 needs hardened builds, with runs isolated from each other and signing material out of reach of user-defined build steps.
SLSA v1.2 has a Build track and a Source track. The weak pattern is a pipeline where any job step can reach the signing key or the OIDC token, so a compromised step can sign whatever it produced. On GitHub, artifact attestations give Build L2 by default; running the build and attestation in a shared, reviewed reusable workflow isolates it from the calling workflow, which is how GitHub describes reaching Build L3. Give pull-request builds no signing identity at all, and tell consumers which builder identity to accept. Provenance proves where and how an artifact was built, not that the code is safe, so SBOMs, scanning and review still apply.
# consumers pin the reusable workflow that is allowed to sign gh attestation verify oci://registry.example/app@sha256:... \ --owner acme \ --signer-workflow acme/build-workflows/.github/workflows/release.yml
Interviewer often follows with: Why does provenance generated by a step inside the same job prove so little?
Link to this questionTag images by commit SHA or digest, retain them per policy — say 90 days for feature builds, longer for release tags — and keep SBOM and provenance beside the artifact so you can prove what shipped.
docker push registry/app:$CI_COMMIT_SHA docker tag registry/app:$CI_COMMIT_SHA registry/app:release-2026-07-24 # registry GC policy: keep digests referenced by release-*Link to this question
I'd pin submodules and dependencies by commit SHA with signature or hash verification, scan diffs in CI, revoke any builds that included the bad commit, and rotate credentials those builds could have seen.
Submodule and go.mod/npm supply-chain attacks succeed when CI floats on a branch ref. Controls: pin SHAs, verify checksums and lockfiles, disable automatic submodule update on untrusted PRs, CODEOWNERS on dependency manifests, and SBOM comparison between builds. Containment: identify digests built after the poison commit, block those digests, rotate CI OIDC roles and any secrets available to the job, audit git history. Long-term: vendoring policy plus Dependabot with human review. Talk pins, attestations, and credential blast radius together.
git submodule status # expect detached SHA, not branch # CI: fail if submodule points to unexpected SHA cosign verify-attestation --type slsaprovenance1 --certificate-identity-regexp "$BUILDER_ID" --certificate-oidc-issuer "$ISSUER" registry/app@$DIGEST
Interviewer often follows with: How do lockfiles fail to protect you if CI runs npm install without ci mode?
Link to this questionReal-world scenarios
I'd freeze prod deploys, roll traffic to the last known-good digest, revoke and rotate the token and any secrets it could reach, then replace long-lived deploy tokens with OIDC and protected environments so a leak can't ship alone.
Contain first: pause pipelines, revoke the PAT or deploy key, invalidate registry push credentials, and verify no second backdoor job was added. Evidence: who used the token, which digests landed, CloudTrail for the window. Remediation: short-lived OIDC, environment approvals, signed images with admission verify, CODEOWNERS on workflow files. Rehearse token-leak → deploy-freeze as a game day. Map the blast radius of every secret that token could read, not only the deploy step.
# revoke token in IdP; disable prod environment kubectl set image deploy/web web=registry/app@sha256:GOOD cosign verify --certificate-identity-regexp '...' --certificate-oidc-issuer '...' registry/app@sha256:GOOD
Interviewer often follows with: How do you prove the rolled-back digest is the same bits that passed staging?
Link to this questionI'd quarantine flakes into a tracked job with ownership and SLOs, keep critical-path tests hard-failing, and ban blanket allow_failure on anything that gates correctness.
allow_failure on the whole suite hides signal. Split must-pass — unit, contract, smoke — from a flake farm with retries, ownership labels, and burn-down. Track flake rate as a KPI; auto-open issues when a test flakes N times. Don't use retries to paper over non-deterministic prod bugs. Gate merge on must-pass only; run broader suites on main or nightly. Distinguish intermittent infra from product non-determinism.
unit:
script: [npm test -- --group=must]
# no allow_failure
flake-quarantine:
script: [npm test -- --group=flaky]
allow_failure: true
retry: { max: 2 }Interviewer often follows with: When is retrying a failed test the wrong mitigation?
Link to this questionWe had a confused-deputy path: deploy trusted a mutable name instead of a digest bound to the trusted build identity. I'd pin by digest, verify signature or provenance, and stop reading latest from shared storage.
Confused deputy: a privileged deployer acts on an attacker-controlled name — latest tag, mutable S3 key, shared cache. Fix: content-address digests, sign at build, verify before promote, separate trust domains so PR builds can't overwrite release artifacts. Use distinct bucket prefixes or repositories per trust level. Admission or deploy-time cosign verify closes the gap if registry tags move. Talk attestation identity — who built — plus immutability — what bits.
DIGEST=$(jq -r .digest build-meta.json) cosign verify --certificate-identity-regexp "$BUILDER_ID" \ --certificate-oidc-issuer "$ISSUER" registry/app@$DIGEST helm upgrade web ./chart --set image.digest=$DIGEST
Interviewer often follows with: Why is passing an artifact URL between jobs still unsafe without a digest?
Link to this questionGitLab removed the CI_JOB_JWT variables in 17.0. I declare an id_tokens entry with the audience the relying party expects, use that token for the Vault or cloud login, and bind the role to claims like project_path, ref, and ref_protected rather than the whole instance.
The old JWTs were exposed in every job; ID tokens exist only in jobs that declare them, and each carries its own aud, so a token minted for Vault is useless against a cloud STS. On the relying side, pin roles to project_path plus a protected ref or protected environment, and give production its own role. On Premium and Ultimate the secrets keyword can fetch Vault values with the ID token directly. Keep protected branches and protected environments in front of any job that can mint a production-audience token, and prove the fix with a feature-branch job that must be denied.
deploy:
id_tokens:
VAULT_ID_TOKEN:
aud: https://vault.example.com
secrets:
DB_PASSWORD:
vault: production/db/password@ops
token: $VAULT_ID_TOKEN
script: ./deploy.shInterviewer often follows with: Why is binding a role to ref alone weaker than binding to ref_protected or a protected environment?
Link to this questionThe analysis was too short and too narrow — it missed diurnal traffic and business KPIs. I'd lengthen windows, require multi-signal abort rules, and block promote until soak SLOs hold under real load.
Fast canaries pass on synthetic or low-traffic periods. Require error rate, latency, saturation, and at least one business SLI; minimum soak under peak-like load or scheduled windows; manual gate for high-risk. Abort must be automatic and faster than human reaction. Pair with feature flags for dark launch. Progressive delivery is an observability contract, not just a YAML weight schedule.
analysis:
interval: 5m
iterations: 12 # ≥1h soak
thresholds:
- error_rate < 1%
- p99_latency < 300ms
- checkout_success > 99%Interviewer often follows with: Which metric would you refuse to omit for a payments API canary?
Link to this questionI'd treat the key as compromised: rotate or disable it immediately, scrub or restrict log retention access, hunt for use in the leak window, and fix the job to mask secrets and ban echo of credential env vars.
Assume public exposure if logs are readable by many engineers or retained in artifacts. Rotate first, investigate second. Enable secret scanning on logs where available; mark CI variables masked; use OIDC so there's no long-lived key to print. Add a CI lint rule that fails on echo or print of known secret env names. Post-incident: who had log read, and were forks able to see it?
# GitLab: masked + protected variable # GitHub: secrets.* never echo; prefer OIDC # fail CI if script matches echo $.*SECRET
Interviewer often follows with: How do you rotate if the printed credential was an OIDC-assumed role session?
Link to this questionSecrets or privileged runners were reachable from a fork PR: a poisoned pipeline execution, or pwn request. I'd isolate fork PRs to untrusted ephemeral runners with no secrets, require approval for first-time contributors, and reserve secrets and OIDC deploy roles for protected refs after merge.
Classic pwn-request: pull_request_target running in the base-repo context while checking out PR code, or shared self-hosted runners with org secrets. GitHub's hardening guide says self-hosted runners should almost never serve public repositories, because any fork PR can run code on them and a non-ephemeral runner can stay compromised for later trusted jobs. Controls: Environment secrets only on main, require approval for first contribution, never checkout PR code in a context that has write tokens, and label runners so fork jobs can't schedule onto privileged pools. Condition OIDC trust on repo plus ref so PR events can't assume the deploy role, and put workflow files under CODEOWNERS. PR build artifacts can be unsigned and discarded; signed provenance comes only from the trusted builder after merge. Cache poisoning or script injection turns one bad run into a full compromise, so scope caches by repo and ref trust. Draw the trust boundary between contributor code and credentialed build.
# secrets: none on pull_request from forks # self-hosted labels: [trusted] only on push to main # OIDC trust sub: repo:org/app:ref:refs/heads/main (never PR events) # workflow_run or environment protection for deploy
Interviewer often follows with: Why is pull_request_target dangerous when combined with checkout of the PR ref?
Link to this questionI'd stop further deploys, restore schema from backup or a forward-fix migration that recreates the column, then redeploy a build compatible with the restored schema — and ban destructive DDL in the same release as the app that needs it.
Rollback of code isn't rollback of schema. Expand/contract discipline: never drop or rename in the same release that removes readers. Recovery options: PITR, replay a dual-write period, or an emergency additive migration. Pipeline should run migrate as a gated job with dry-run and require backward-compatible migrations for rollback windows. Game-day this failure. Rollback SLA must include the data plane, not only image tags.
# restore column (forward fix) or PITR to pre-migrate # deploy last app digest that matches restored schema # block prune/drop migrations without a compatibility window
Interviewer often follows with: How would expand/contract have prevented this outage?
Link to this questionI'd explain build-once/promote-many: prod should track the digest that was tested, not a tag rebuilt later. Compare attestations, freeze floating tags, and show the signed provenance of the running digest.
Rebuilding invalidates the evidence chain — dependency drift, base image updates, non-hermetic builds. Controls: immutable digests in prod manifests, retention of release artifacts, cosign/SBOM attached at original build, and CI that refuses to retag release digests. If rebuild is required for a CVE rebase, treat it as a new release with full retest. Reproducibility and immutability both matter, but prod pins bits.
kubectl get deploy web -o jsonpath='{.spec.template.spec.containers[0].image}'
cosign verify --certificate-identity-regexp "$BUILDER_ID" --certificate-oidc-issuer "$ISSUER" registry/app@sha256:PROD
# do not docker build -t v1.2.3 again and pushInterviewer often follows with: What breaks if your registry allows tag overwrite on v1.2.3?
Link to this questionI'd make the required status check the actual test job — or a gate that ANDs all must-pass jobs — and never let an artifact-upload or notify job be the merge gate.
Branch protection that requires a misnamed or soft job is a common footgun. Use a single required CI-passed job that needs unit, integration, and security — or GitLab's equivalent. Disable continue-on-error on required paths. Dashboards that show only the last job mislead on-call. Treat required checks as a security control and review them like IAM.
gate: needs: [unit, integration, gitleaks] script: [echo ok] # branch protection: require "gate", not "upload-coverage"
Interviewer often follows with: How do you handle a known-flaky optional job without making the gate optional?
Link to this questionI'd introduce a severity threshold in warn mode with a tracked exception list, fix or rebase the worst images first, then flip to fail-closed for criticals on main while leaving advisories non-blocking.
Big-bang blocking breaks trust in security. Pattern: inventory top base images, set grace windows per severity, CODEOWNERS exceptions with expiry, and alternate patched bases. Fail CI on critical/high for release branches first. Pair with digest pins and rebuild pipelines. Measure MTTR for CVE gates. Risk-based enforcement plus exception hygiene beats perpetual warn-only.
trivy image --severity CRITICAL --exit-code 1 app:$SHA # exceptions: .trivyignore with ticket + expiry date
Interviewer often follows with: How do you prevent .trivyignore from becoming permanent debt?
Link to this questionRelated
- Cheat sheetGitLab CI/CD cheat sheet
- Cheat sheetGit cheat sheet
- Interview guideDevSecOps & supply-chain security interview questions
- CourseSecure CI/CD with GitLab
- CourseSoftware supply chain security
- CourseSoftware supply chain in depth
- Field noteStop leaking secrets in CI logs: masking and OIDC
- Field noteHardening self-hosted GitLab runners
- Field noteAttesting builds with SLSA provenance in CI
Primary references
Found a technical issue on this page? Report it with the tool version you used and the behavior you saw. How resources are maintained.