DevSecOps & supply-chain security interview questions
Practice DevSecOps interview answers from shift-left scanning through supply-chain signing, runtime detection, and security culture.
Levels run Beginner → Intermediate → Advanced → Expert. Answers are phrased the way you would say them in an interview; Advanced and Expert answers add the deeper reasoning, a diagram where it helps, and the follow-up an interviewer often asks next.
Fundamentals
It's moving security earlier into design and development — threat modeling, linting, and CI scanning — so issues are cheap to fix, instead of waiting for a late pen-test or a production incident.
design → threat model code → SAST / secret scan / lint CI → SCA, image scan, IaC, sign prod → admission, runtime, monitorLink to this question
SAST looks at source or bytecode without running it. DAST attacks a running app from the outside. SCA finds known CVEs in third-party deps. IAST instruments the running app to watch real code paths. They catch complementary classes of bugs — I wouldn't pick just one.
SAST → Semgrep, CodeQL (your code) SCA → npm audit, pip-audit, Trivy fs DAST → ZAP / Burp against a URL IAST → agent inside the app runtimeLink to this question
A CVE IDs a public vulnerability; CVSS scores it 0–10 on exploitability and impact. I wouldn't prioritize on score alone — I'd weigh reachability, exposure, and exploit signals like EPSS and CISA KEV. An unreachable 9.8 can wait longer than a reachable 6.5 on a public login path.
1) in CISA KEV or high EPSS? 2) reachable from your code / internet-facing? 3) fixable version available? 4) then CVSS (v4.0 where published, else v3.1) as a tie-breakerLink to this question
It inventories what is installed, OS packages from the package manager's database (dpkg, rpm, apk) and language packages from lockfiles and manifests, then matches each name and version against advisory data. For distro packages the reliable source is the distro's own advisories, because distros backport fixes without changing the upstream version number.
Trivy, for example, uses vendor data for OS packages (Debian's security tracker, Alpine's secdb, Red Hat's OVAL data) and the GitHub Advisory Database plus language-specific feeds for application dependencies. Matching on the upstream version alone would flag a backported package: Trivy's documentation cites CVE-2023-32681, which affects upstream requests 2.20.0 but is fixed in Red Hat's 2.20.0-3 package. Two scanners disagree because they read different advisory sources, rate severity differently (vendor versus NVD), detect different packages (a binary with no package metadata can be invisible to one of them), and refresh their databases at different times. So I pin the scanner version, keep the scan date with each result, and triage a finding against the advisory for the package I actually ship.
trivy image --format json registry/app@sha256:... > scan.json # every package found, plus its matches trivy image --severity HIGH,CRITICAL --ignore-unfixed registry/app@sha256:...
Interviewer often follows with: A scanner reports a critical CVE in openssl on your Debian image, but Debian lists it as fixed in the version you run. Which do you believe, and why?
Link to this questionEvery identity — human, service, pipeline — gets only the permissions it needs for the shortest time. That limits blast radius when something is compromised, and it's the foundation under zero-trust and supply-chain defense.
# Prefer OIDC short-lived cloud roles scoped to repo+branch # over long-lived access keys stored as CI secretsLink to this question
It's systematically asking what can go wrong before you build: draw trust boundaries and data flows, brainstorm threats (e.g. STRIDE), and decide mitigations. It catches design flaws scanners will never see.
# 60-minute PRD review: # - what's sensitive? who can call what? # - spoofing / tampering / repudiation / info disclosure / DoS / elevation # - ticket mitigations before coding startsLink to this question
I'd say DevSecOps embeds automated controls and shared ownership into the delivery path so security is continuous feedback, not a release-week veto. Security still sets the bar; engineering owns fixing inside the sprint.
golden pipeline template: secret-scan → sast → sca → build → image-scan → sign → deployLink to this question
Pipeline security & scanning
I'd run Trivy/Grype against the built image (and continuously in the registry), failing on fixable high/criticals. Rescan matters — new CVEs appear for images that have not changed.
trivy image --severity HIGH,CRITICAL --ignore-unfixed --exit-code 1 app:1Link to this question
I'd use lockfile-aware scanners in CI (pip-audit, npm audit, Trivy fs, Grype) so vulnerable transitive deps fail the build. The hard part is a failure policy the team keeps enabled — not installing another tool.
pip-audit -r requirements.txt npm audit --audit-level=high trivy fs --scanners vuln .Link to this question
I'd run static scanners — Checkov, Trivy config (tfsec's successor), KICS — and Conftest on Terraform plans and K8s YAML so public buckets, open security groups, and privileged Pods fail in the pipeline before apply.
checkov -d . --compact trivy config . terraform show -json plan.out | conftest test -Link to this question
I'd block merges with pre-commit plus CI scanners like gitleaks or trufflehog. On a hit, rotate the credential immediately — rewriting history isn't enough — and move secrets into a manager.
gitleaks git --redact . # HIT → rotate in IdP/cloud NOW → revoke old → store in Vault/SMLink to this question
I'd narrow to high-confidence rules, baseline existing debt and block only on new actionable findings versus main, add suppressions with expiry and owner, and measure reopen/disable rates. A quieter gate that stays on beats a loud one that's bypassed.
Noise kills shift-left. Start with curated rule packs (OWASP top classes you actually see), exclude generated code, and require every suppression to have reason + expiry in Git. Separate “break build” from “report only” severities, and block PRs only on new high/critical findings compared with main (differential scanning); burn the existing backlog down as a tracked epic, not a merge blocker. Publish mean-time-to-fix and false-positive rate as team metrics. Pair SAST with codeowners so the right people see findings. Re-enable gradually: one rule family at a time. Security owns scanner quality jointly with platform. The goal is sustained adoption, not scanner marketing counts.
# .semgrepignore / inline nosemgrep with ticket + expiry # CI fails only on rules tagged confidence:high severity:HIGH or CRITICAL (legacy ERROR maps to HIGH) # PR gate: NEW findings vs main only; backlog = weekly burn-down, not a merge gate
Interviewer often follows with: How would you prove the gate still catches real bugs after quieting it?
Link to this questionI'd fail on fixable, reachable, high-confidence criticals; route the rest to dashboards; require expiring justified exceptions; and give fast local/CI feedback. Then ratchet severity up as noise drops.
Policy design is basically product management — SLOs for scan duration, clear ownership of CVE backlog, and exception workflows that security can audit. Prefer deny-lists of known-bad packages and KEV-driven emergency fails. For containers, ignore-unfixed avoids impossible work while you track base-image refresh. I never make 'critical CVSS' alone the only gate without reachability/KEV context or you train people to ignore red builds.
# week 1: fail KEV + critical fixable in direct deps # week 4: add transitive highs with available patches # exceptions.yaml: cve, owner, expiresOn, compensating_control
Interviewer often follows with: What compensating controls make a temporary CVE exception acceptable?
Link to this questionI'd run lightweight smoke DAST on every deploy to staging; run deeper authenticated scans on a schedule or pre-release. Fail the release on high-confidence findings with known exploit paths; don't block every PR on a 40-minute crawl.
PR: Semgrep + SCA (minutes) staging deploy: ZAP baseline against /health and login nightly: full authenticated ZAP/Burp suite → ticketsLink to this question
I'd put secret scan, SAST, SCA, build, image scan, sign, and deploy gates in a reusable template every team inherits — so the secure path is the default path, not a wiki they skip.
secret-scan → sast → sca → test → build → image-scan → sbom → sign → deployLink to this question
I'd treat runners as production: ephemeral VMs/containers, no long-lived cloud keys, isolated networks, patched images, restricted who can run privileged workflows, and never reuse the same runner pool for untrusted fork PRs and prod deploys.
Poisoned pipeline execution thrives on shared, over-privileged runners. Separate trust tiers: untrusted PR runners without secrets; trusted release runners with OIDC and signing. Disable pull_request_target anti-patterns, pin actions by SHA, and monitor runner egress. Hardened runners are as important as app SAST — a compromised runner can mint signed “trusted” artifacts if signing keys/OIDC roles are reachable.
fork PR jobs: no secrets, network-limited runner main release: OIDC role + cosign, no fork code execution
Interviewer often follows with: What is a pwn request / poisoned pipeline execution?
Link to this questionNew CVEs land against unchanged images. I'd rescan the registry continuously and rebuild or rebase when fixable criticals appear — a one-shot CI scan isn't enough.
trivy image --exit-code 1 registry/app@sha256:abc # cron / registry scanner over all prod tagsLink to this question
Patch with parameterized queries, add a regression test or scanner regression case, deploy behind staging DAST smoke, and consider a WAF rule as compensating control while you harden related queries.
Legacy constraints tempt “just suppress.” Prefer fix + minimal characterization test. If a full rewrite is impossible, mitigate with input validation at the edge, least-privilege DB creds, and monitoring for SQLi patterns. Track as risk acceptance only if fix is truly deferred with expiry. The aim is layered mitigation and a path to delete the waiver.
# bad: query = "SELECT * FROM u WHERE id=" + userInput # good: parameterized statement / ORM binder # add Semgrep regression + staging ZAP smoke on the route
Interviewer often follows with: When is a WAF rule an acceptable long-term control versus a bridge?
Link to this questionSBOM, signing & provenance
An SBOM is a Software Bill of Materials — components and versions in an artifact, usually SPDX or CycloneDX. I generate it from the built image so when the next Log4Shell drops I can answer 'are we affected?' in minutes.
syft app:1 -o cyclonedx-json > sbom.json grype sbom:sbom.json --fail-on highLink to this question
Either is defensible; SPDX is ISO/IEC 5962 and CycloneDX is ECMA-424. I pick what my consumers and tools ingest, then spend the effort on completeness: transitive dependencies, PURLs, generated from the built artifact, stored beside the digest.
syft app@sha256:abc… -o cyclonedx-json > sbom.cdx.json syft app@sha256:abc… -o spdx-json > sbom.spdx.json # latest specs: CycloneDX 1.7, SPDX 3.0 # check which spec version your generator actually emits
Interviewer often follows with: What would you check before accepting a supplier's SBOM as evidence?
Link to this questionI'd sign in CI with cosign (preferably keyless OIDC identity), store signatures/attestations in the registry, verify again in the deploy pipeline, and enforce at admission so unsigned images can't schedule even if someone bypasses CI.
Keyless signing ties the signature to the workflow identity (repo, ref, actor) via a short-lived cert — no long-lived cosign keys to leak. Pin certificate-identity / OIDC issuer in verify policies. Attest SBOMs and provenance (SLSA) alongside the image. Admission (Kyverno ImageValidatingPolicy, which replaces the deprecated ClusterPolicy verifyImages rules, or Sigstore policy-controller) is the hard gate; CI verify is the fast feedback. Rotate trust by updating policy allowlists, not by hoping old keys disappear. Mutable tags undermine signing — promote digests.
cosign sign app@sha256:abc… # keyless in GitHub Actions OIDC cosign verify app@sha256:abc… \ --certificate-identity https://github.com/acme/app/.github/workflows/release.yml@refs/heads/main \ --certificate-oidc-issuer https://token.actions.githubusercontent.com
Interviewer often follows with: For keyless signatures, what exactly do you pin in the verify policy, and why is the issuer alone not enough?
Link to this questionProvenance is signed metadata about how, from what source, and by which builder an artifact was produced. SLSA tiers raise integrity requirements so consumers can trust origin — higher levels need hardened builders that can't forge their own provenance.
Provenance answers “was this built from commit X by builder Y I trust?” SLSA Build track (spec v1.2): L1 provenance exists → L2 a hosted build platform signs it → L3 a hardened platform isolates builds and keeps signing secrets away from build steps, so forging provenance takes an exploit. That defeats a class of compromised CI steps that produce “clean” looking artifacts. Verification belongs in admission/CD: reject missing or mismatched provenance. Pair with hermetic builds and locked dependencies so the provenance is meaningful.
cosign verify-attestation --type slsaprovenance1 \ --certificate-identity https://github.com/acme/app/.github/workflows/release.yml@refs/heads/main \ --certificate-oidc-issuer https://token.actions.githubusercontent.com \ app@sha256:abc… # policy: builder identity must match org's GitHub Actions workflow
Interviewer often follows with: What stops a malicious workflow in a forked PR from producing “valid” keyless signatures?
Link to this questionAttackers target build systems and maintainer trust, not only your application code. I'd defend with hermetic, verifiable builds, signed provenance, least-privilege CI, and scrutiny of dependency maintainership — not just CVE scanning.
SolarWinds showed a poisoned pipeline can ship trusted updates; xz showed social engineering of maintainership. Scanners wouldn't have saved you if the malicious code was the intended release. Controls: isolated builders, two-party review on release workflows, pin actions by SHA, disable privileged fork workflows, reproducible builds where feasible, and monitor for anomalous maintainer changes on critical deps. SBOMs help response; provenance + identity help prevention.
pin actions to SHA OIDC to cloud (no long-lived keys) no secrets on pull_request from forks sign + attest + admit-time verify
Interviewer often follows with: How would you detect a malicious GitHub Action behaving like a clean build?
Link to this questionI'd pin versions and hashes via lockfiles, use a private proxy with allow-lists so internal names can't be hijacked from the public registry, vet new packages, and prefer few well-maintained dependencies.
# internal scope @acme/* only from private registry # public packages via pull-through with firewall rules # lockfile + npm ci / pip hash checking in CILink to this question
I publish a VEX statement, OpenVEX or CycloneDX VEX, marking the CVE not_affected for that image digest with a justification such as vulnerable_code_not_in_execute_path. I sign and attach it beside the image and feed it to the scanner, so the decision is reviewable and owned instead of living in a local ignore file.
vexctl create --product="pkg:oci/app@sha256%3Aabc…" \ --vuln="CVE-2026-1234" --status="not_affected" \ --justification="vulnerable_code_not_in_execute_path" > app.vex.json trivy image --vex app.vex.json registry/app@sha256:abc… # Trivy marks VEX support experimental
Interviewer often follows with: What should happen to a not_affected statement when a code change starts calling the vulnerable function?
Link to this questionScanning looks for known vulns in the contents. Signing proves who built it and that it hasn't been tampered with. I'd do both — a signed image can still be full of CVEs, and a clean scan doesn't prove provenance.
trivy image app@sha256:abc # content risk cosign verify app@sha256:abc # origin/integrity (+ --certificate-identity / --certificate-oidc-issuer)Link to this question
Query stored SBOMs/attestations for the affected package/version, page owners of hits on internet-facing tiers, patch or mitigate, and verify with a fresh SBOM after rebuild — hours, not days of grepping repos.
Prerequisite is generating and indexing SBOMs at build (Syft/Trivy) and retaining them beside digests. Response playbook: identify CPE/PURL, search inventory, classify exposure, ship fixes via golden base images where possible, and communicate customer impact. Without SBOMs you rediscover ownership manually. Pair with admission that prefers digests you have SBOMs for. This is a classic senior supply-chain question after Log4Shell.
grype sbom:sbom.json --fail-on critical # or central inventory: SELECT service WHERE pkg=log4j AND ver < fixed
Interviewer often follows with: Where do you store SBOMs so incident responders can find them at 2am?
Link to this questionSignature proves who signed with a key/identity, not that the build was hermetic — I require provenance/SLSA attestations bound to a trusted builder identity and verify those before deploy.
Keyful signing on a shared mutable runner is weak: malware in the builder signs malware. Mitigations: hermetic/reproducible builds, isolated builders, provenance attestations (SLSA), and policy that checks builder identity (the Fulcio certificate's GitHub workflow identity), not only “sig exists.” Compare SBOM to source. Rotate if builder compromise is suspected. Keep integrity (the signature) separate from build trustworthiness (provenance).
cosign verify-attestation --type slsaprovenance1 \ --certificate-identity "$BUILDER_WORKFLOW_URI" \ --certificate-oidc-issuer https://token.actions.githubusercontent.com \ registry/app@$DIGEST # policy: builder identity must be repo/workflow allow-list
Interviewer often follows with: What changes between SLSA L2 and L3 that matters here?
Link to this questionRuntime & architecture
I'd expect runtime detection like Falco (eBPF syscall rules) to alert on shell spawn / unexpected egress, plus NetworkPolicy egress defaults, read-only rootfs, and dropped capabilities so the blast radius is small while humans respond.
Build-time scanning can't see post-exploit behavior. Falco/eBPF watches syscalls: shell in container, write to /etc, unexpected outbound DNS. Tune rules to reduce noise; route high-severity to page, rest to SIEM. Pair with PSS restricted, no privilege escalation, and egress NetworkPolicies. Response: isolate the Pod/node, snapshot for forensics, rotate credentials the workload could reach, and patch the entry point. Admission and signing still matter — runtime is the last layer, not the first.
- rule: Shell in a container desc: shell spawned inside a container condition: spawned_process and container and proc.name in (bash, sh) output: "shell in container (pod=%k8s.pod.name)" priority: WARNING
Interviewer often follows with: How would you distinguish a legit kubectl exec debug session from an attacker shell?
Link to this questionNo trust from network location alone — every request and every pipeline step is authenticated, authorized, and least-privilege. Identity (workload + human) is the control plane; mTLS, short-lived creds, and continuous verification replace flat VPN trust.
Concretely for delivery: OIDC from CI to cloud (no static keys), signed artifacts verified at admission, SPIFFE/mTLS between services, per-request authz (OPA), and segmented networks as a backstop not the only control. Developers access prod through audited break-glass, not standing admin kubeconfig. Zero trust is a program of identities and policies — buying a vendor alone doesn't implement it. Tie to DevSecOps by making the paved road issue identities and verify signatures automatically.
CI OIDC → cloud role (repo/branch conditioned) cosign sign → registry admission verifyImages → schedule runtime mTLS + Falco → detect
Interviewer often follows with: Where do people fake zero trust while still using long-lived cluster-admin kubeconfigs?
Link to this questionLayer controls so one failure isn't fatal: hardened non-root image, SBOM+scan+sign in CI, PSS/admission policy, least-privilege RBAC and NetworkPolicy, secrets from a manager, runtime detection, and monitored SLOs — assume any single layer can fail.
Map layers to kill chain stages: prevent bad code (SAST/SCA), prevent bad artifacts (sign/provenance), prevent bad schedules (admission/PSS), limit blast radius (RBAC, NetPol, mTLS), detect escape (Falco), respond (runbooks, revoke). GitOps keeps the desired hardened state honest. Name the layers and what each one stops, not a list of tools. Call out etcd encryption and backup as part of the secrets/data layer.
image (non-root, minimal) CI (SCA/SAST/scan/sign) admit (PSS + verifyImages) identity (SA RBAC + NetPol) runtime (Falco) + respond
Interviewer often follows with: Which layer do you invest in first for a greenfield team with three engineers?
Link to this questionOnly for narrow, high-confidence behaviour where a false positive is cheaper than a missed attack, like a shell or package manager exec in a locked-down prod namespace or writes to sensitive host paths. Everything else stays detect-and-alert until the rule has a measured precision record.
Tetragon applies TracingPolicy filters in eBPF and can act before a syscall completes: an Override action makes the call return an error, and a Signal action can send SIGKILL. Its docs say to pair them, because SIGKILL alone may not stop an operation already in progress. Falco detects and alerts; response is handled by tools downstream of it. Roll out enforcement the way you roll out admission policy: observe first, scope by namespace or workload labels, keep a break-glass path, and alert on every enforced action so a bad rule shows up as an incident instead of silent pod crashes.
observe: TracingPolicy without actions, two weeks scope: prod-* namespaces, app workloads only enforce: Override (+ Signal) on the narrow match alert: every enforced event → owning team
Interviewer often follows with: What happens to availability if an enforcement rule matches a legitimate init script after a base image update?
Link to this questionI isolate the node/workload (cordon/drain, network deny), preserve forensic evidence, rotate secrets the pod could reach, and patch the escape path — not reboot-wipe before capture if evidence matters.
Order: detect (Falco/runtime), contain (cordon, NetworkPolicy/cloud SG, pause scheduling), eradicate (kill malicious pods, patch CVE/misconfig like privileged+hostPath), recover (replace node), lessons. Assume cloud credentials and SA tokens are burned. Snapshot disks if legal or forensics require it. Containment comes before curiosity, and evidence before rebuild when it is needed.
kubectl cordon <node> kubectl drain <node> --ignore-daemonsets --delete-emptydir-data # revoke IRSA/role sessions the pods used; rotate Vault leases
Interviewer often follows with: Which Kubernetes privileges most often enable escape to the host?
Link to this questionVulnerability risk & culture
Rank by real risk: KEV/EPSS, reachability, internet exposure, asset criticality, and fix availability — not raw CVSS order. Fix reachable exploited issues on crown-jewel services first; accept or defer the rest with recorded rationale.
Backlogs explode when every scanner finding is equal. Build a risk queue: (1) KEV present (2) public-facing + high EPSS (3) reachable in your call graph (4) privileged context (5) everything else. Automate reachability where tools allow; otherwise sample by service tier. Time-box “fix all criticals” mandates — they create exception theater. Report MTTR for KEV items as the executive metric. Risk acceptance is explicit, owned, and expired — not silent ignore.
priority=P0: KEV or known wormable + exposed P1: reachable high on tier-0 services P2: fixable highs with patch, scheduled P3: accepted with owner + review date
Interviewer often follows with: How would you handle a critical CVE in a base image you don't control yet?
Link to this questionA deliberate, time-bounded decision to run with a known risk, documented with owner, rationale, compensating controls, and expiry — reviewed, not ignored. It's a governance artifact, not a way to silence scanners forever.
Good acceptance tickets include: asset, vulnerability, business justification, compensating controls (WAF, disable feature, network isolate), residual risk, expiry, and approving authority. CI suppressions should link to that ticket and fail when expired. Security aggregates acceptances for audit. A permanent .trivyignore with no owner is the anti-pattern. Tie acceptance to error budgets and product risk, not to security theater.
cve: CVE-2026-1234 service: billing-api owner: billing-oncall compensating: not reachable; egress denied expiresOn: 2026-10-01 ticket: SEC-4412
Interviewer often follows with: Who should be allowed to approve production risk acceptance in your org?
Link to this questionChampions are engineers in each team with extra training and a direct line to security — they review designs, triage scanner noise, and spread paved-road patterns. Scale comes from multipliers and golden paths, not central ticket bottlenecks.
Champions need time allocation, a community (office hours, Slack), and authority to block clearly dangerous patterns. Feed them threat-model templates, secure defaults (hardened base images, pipeline templates), and recognition. Measure champion-led fixes and reduction in repeated finding classes. Avoid making champions unpaid gatekeepers for every PR — focus on high-risk changes and mentoring. Pair with platform guardrails so the default path is already secure.
monthly: top CVE classes + one threat-model demo per team: own suppressions + secure pipeline template security: office hours + escalate path for P0
Interviewer often follows with: How would you keep champions from burning out or becoming shadow security reviewers for everything?
Link to this questionEPSS estimates the probability a CVE will be exploited in the wild in the next 30 days. I use it with CVSS and KEV — a medium CVSS with high EPSS on a reachable path often jumps the queue ahead of an unreachable critical.
# Priority bump if: # - CISA KEV listed, OR # - EPSS above your threshold (e.g. 0.5) AND reachableLink to this question
Cap open exceptions, require expiry and compensating controls, report exception age to leadership, and auto-fail expired waivers in CI. Measure KEV MTTR, not waiver count.
Unlimited waivers recreate the pre-DevSecOps status quo. Governance: max active waivers per service tier, monthly review board, and dashboards of waiver debt. Pair with investment in base-image maintenance so teams aren't forced to waive unfixable noise. Culturally, praise closed KEVs and reduced mean age of criticals. In an expert interview I want metrics + incentives, not another scanner.
# parser fails build if exceptions.yaml expiresOn < today # report: P95 age of open critical exceptions
Interviewer often follows with: How would you handle a vendor dependency with no patch for 180 days?
Link to this questionI patch the golden base once, rebuild dependents in waves by exposure tier, block new deploys of vulnerable digests at admission, and track KEV MTTR as the KPI.
Rebuilding 200 repos ad-hoc fails. Platform move: rebuild and resign the golden base, auto-PRs or image automation for consumers, prioritize internet-facing and data-tier services, and use admission/cosign policies to refuse old bases after a deadline. Communicate freeze windows. Offer a short exception with expiry for non-exposed internal tools. The shape of the answer: golden image, admission deadline, tiered waves.
# after T+72h: verifyImages / policy denies bases older than patched digest # dashboard: % services on patched base
Interviewer often follows with: How would you handle a service that can't rebuild because of a broken upstream?
Link to this questionI refuse silent merge: require a time-boxed risk acceptance with owner, compensating controls, monitoring, and a tracked removal ticket — or offer a safer alternative (custom CA, mTLS) that unblocks the feature.
DevSecOps is escalation and design, not only scanners. Document blast radius (MITM, credential theft), insist on expiry, and add detection (traffic to insecure endpoints). Prefer fixing trust stores. If leadership accepts risk, record it in the risk register. In an expert interview I want persuasion + alternatives + expiry, not pure veto theater.
risk: disable TLS verify to legacy vendor owner: svc-team lead expires: 2026-08-24 compensating: private network + allowlist + alert on dest exit: vendor cert fix / custom CA
Interviewer often follows with: What compensating control is insufficient for disabling TLS verify on the public internet?
Link to this questionReal-world scenarios
I triage by reachability, exposure, and KEV/EPSS — then either mitigate (WAF, remove package, alternate base), accept risk with expiry, or rebuild when a fix exists — never “ignore forever because unscored.”
No-fix criticals are common. Is the package loaded and reachable? Internet-facing? In CISA KEV? Can you drop the component or switch distros? Compensating controls buy time; admission can still block if exploitability is high. Document owner, expiry, and re-scan trigger when a fixed version appears. Communicate product risk clearly — CVSS alone isn't the answer.
CVE-2026-XXXX critical no-fix reachable: yes (httpd module loaded) KEV/EPSS: elevated action: migrate to distroless/alternate base OR risk-accept until T+14 ticket: SEC-9921 expires 2026-08-07
Interviewer often follows with: What changes if the same CVE is in an unreachable build-time-only tool deleted from the final image?
Link to this questionI switch to lockfile-aware SBOM generation that walks the full graph (Syft/Trivy/cdxgen on the built artifact), fail CI when the SBOM is incomplete, and scan that SBOM — not a hand-written package list.
Shallow SBOMs create false confidence. Generate from the built image or language lockfiles after resolve, emit CycloneDX/SPDX with transitive nodes, and feed the same document to SCA and admission. Attest the SBOM beside the image digest. Periodically diff “SBOM packages” vs “scanner findings” to catch generator gaps. SBOM quality is a control, not a checkbox file.
syft scan registry/app@$DIGEST -o cyclonedx-json > sbom.json grype sbom:sbom.json --fail-on critical cosign attest --predicate sbom.json --type cyclonedx ...
Interviewer often follows with: Why is generating an SBOM only from package.json without the lockfile dangerous?
Link to this questionTags are mutable — verify must pin digest (image@sha256:…) or an immutable digest reference from the deploy manifest, never a floating tag alone.
cosign verify on a tag only checks whatever digest the tag currently points to at pull time; tag move = different bits under the same name. The fix: CI writes digest into GitOps, admission verifies that digest + signature/identity, and registry tag immutability where supported. Prefer keyless identities bound to repo/workflow. Pair with verify-images policies that reject tag-only refs.
# BAD image: registry/app:prod # GOOD image: registry/app@sha256:abc123... cosign verify --certificate-identity-regexp 'https://github.com/org/app/.github/workflows/.*' \ --certificate-oidc-issuer https://token.actions.githubusercontent.com \ registry/app@sha256:abc123...
Interviewer often follows with: How does an attestation (provenance) still help if someone can move tags?
Link to this questionI tune rules and scopes first: suppress known-good builder namespaces, keep prod shell-spawn hot, and scope exec exceptions to break-glass identities with owners and expiry. Then I measure precision before touching paging thresholds.
Alert storms destroy runtime detection. Split environments: noisy rules in build namespaces become metrics-only; prod app namespaces stay high-signal (shell, sensitive mounts, unexpected network). Legitimate prod debugging goes through a ticketed break-glass role or a debug Namespace, preferably with ephemeral containers and a short TTL, so exec-into-prod stays noisy by design for everyone else. Use allowlists for platform DaemonSets and known sidecars. Capacity: sidekick/queue so Falco itself stays healthy. Review top rules and alert precision weekly, and treat exceptions as products with owners, not silenced rules. Deleting the sensor is the failure mode — retuning is the job.
# rule: Spawned shell in container # priority: WARNING in ns=build-* (metrics only) # priority: CRITICAL in ns=prod-* (page) # exception: sa=incident-debug TTL=2h
Interviewer often follows with: What is the risk of a global Falco silence during an incident?
Link to this questionI revoke/rotate trust immediately, block admission of signatures from the bad key/identity, re-sign or rebuild from trusted builders, and hunt for images signed during the exposure window.
Compromise means signatures no longer prove integrity. Steps: revoke key in KMS/Cosign trust root or remove Fulcio/GitHub identity from policy allow-lists; fail closed on verify; invalidate suspicious digests; rotate any secrets that builders held; rebuild critical services on clean runners; forensics on CI logs. Prefer short-lived keyless identities to reduce blast radius next time. Treat signing trust like production credentials: revoke first, explain later.
# policy: remove compromised key/identity from verifyImages allow-list # admission: deny images signed only by revoked key # CI: new keyless identity / rotated KMS key # inventory: images signed between T0 and T_revoke
Interviewer often follows with: Why is “just generate a new key and keep accepting the old one” dangerous?
Link to this questionI tier controls: secrets and critical reachable CVEs stay merge-blocking; medium/low and docs paths become warn/async; trunk gets continuous scanning with SLO-bound debt.
One severity for all paths is why teams mutiny. Design: path filters (docs/, *.md skip SCA), differential vs full scans, severity×reachability matrix, and async tickets for non-blocking debt with age SLAs. Champions help teams fix real issues fast. Leadership dashboard shows blocked-PR time vs true-positive rate. DevSecOps optimizes for risk-reduced throughput, not maximum red X's.
block merge: secret scan hit OR reachable critical CVE warn + ticket: medium SCA, style SAST skip: docs/**, *.md (except secret scan) nightly: full fleet rescan → backlog SLO
Interviewer often follows with: How would you stop warn-only findings from rotting forever?
Link to this questionYes for the build supply chain — I scan and harden builders separately, use ephemeral runners, and make sure build tools can't reach production credentials.
Final-image green doesn't mean CI is safe. Compromised build agents sign and push trusted malware. Controls: patched AMIs/images for runners, SCA on builder images, network egress allowlists, OIDC short-lived cloud roles, no long-lived deploy keys. Separate build and runtime SBOMs. The trust boundary has to include the pipeline itself.
trivy image ci-runner:2026.07 # runner: ephemeral, no prod kubeconfig # deploy via OIDC federated role from trusted workflow only
Interviewer often follows with: What attestation helps prove which builder produced a digest?
Link to this questionI correlate process tree, network destinations, binary hash, and node/pod timeline — confirm or dismiss with evidence, then tune the rule if it was noisy — never silence on vibes alone.
CPU can look “normal” on multi-core nodes while a small miner runs. Check Falco fields (proc, exe, connection), compare image layers to known binaries, inspect DNS/egress, and see if the pod was unexpected. If true: isolate, rotate, replace node if escape suspected. If false: refine rule (known sidecar paths). Feed outcomes back into precision metrics: detection quality is a loop, not a one-shot page.
# Falco: proc.name, fd.sip, container.image.repository kubectl get po -o wide; kubectl describe po # egress: DNS to mining pools? unexpected listen ports? # hash exe vs image layer
Interviewer often follows with: When would you cordon the node even if CPU graphs look flat?
Link to this questionI rotate every exposed secret immediately, rebuild without secrets using BuildKit secret mounts or runtime injection, scrub/republish digests, and add a secret-scan gate that blocks this class of PR.
Earlier layers keep a secret even after a later RUN rm, and ARG/ENV values persist in the final image. Assume compromise: rotate cloud keys, tokens, and DB passwords; deny old digests at admission; force redeploy. Prevention-wise, secret scanning (gitleaks/trufflehog) on PR + image history checks, educative paved-road Dockerfile. Never “it was only staging.”
# 1) rotate credentials in Vault/IdP # 2) rebuild with --secret id=npm,src=... (BuildKit) # 3) cosign sign new digest; admit only new digest # 4) gitleaks git on every PR + gitleaks git --pre-commit --staged locally
Interviewer often follows with: When does a secret used in a builder stage end up in the final image of a multi-stage build, and when does it not?
Link to this questionI enforce verify-images fail-closed with HA webhooks, digest-pinned GitOps, and a ticketed break-glass Namespace or annotation with short TTL and audit — hotfixes still sign via a break-glass signer path.
Availability vs integrity: unsigned emergency images are a conscious exception. The normal path is keyless signing in CI; break-glass signer in a hardware-backed or tightly controlled identity; PolicyException TTL; alert on every exempt admit. Webhook HA so verify outage ≠ silent fail-open. Write the break-glass runbook before the SEV-1.
# Kyverno ImageValidatingPolicy / Sigstore policy-controller: deny unsigned # break-glass: annotation secops.io/unsigned-ok=ticket:INC-1 expires=2h # alert: any admit with that annotation # hotfix: sign with break-glass identity, still prefer signed
Interviewer often follows with: What secondary control still helps if someone bypasses verify during break-glass?
Link to this questionI usually warn rather than hard-block when reachability is high-confidence negative, still track upgrade, and reserve hard-block for KEV/internet-facing or uncertain reachability — and I document the reasoning.
Reachability reduces noise but isn't perfect (reflection, native code, future call paths). Policy: auto-waive only with tool confidence + human override for KEV; require upgrade within an SLO; re-open if code paths change. Prefer upgrading anyway when cost is low. Risk-based policy beats a binary “all criticals fail.”
if KEV or public-facing: block else if reachability=not-callable (high confidence): warn + 30d upgrade SLO else: block critical / high-fixable
Interviewer often follows with: What makes reachability analysis wrong in polyglot or heavily reflective apps?
Link to this questionRelated
- Cheat sheetGitLab CI/CD cheat sheet
- Cheat sheetGit cheat sheet
- Interview guideCI/CD interview questions
- Cheat sheetPython for DevSecOps cheat sheet
- CourseSecure CI/CD with GitLab
- CourseSoftware supply chain security
- CourseSoftware supply chain in depth
- Field noteStop leaking secrets in CI logs: masking and OIDC
- Field noteHardening self-hosted GitLab runners
- Field noteAttesting builds with SLSA provenance in CI
Primary references
Found a technical issue on this page? Report it with the tool version you used and the behavior you saw. How resources are maintained.