Policy-as-code interview questions
Practice policy-as-code interview answers covering Rego, Conftest, Gatekeeper, Kyverno, safe rollouts, and exemption governance.
Levels run Beginner → Intermediate → Advanced → Expert. Answers are phrased the way you would say them in an interview; Advanced and Expert answers add the deeper reasoning, a diagram where it helps, and the follow-up an interviewer often asks next.
Fundamentals
It's writing your security, compliance, and ops rules as versioned, testable code that machines enforce, instead of wiki pages and manual review. Code is reviewed, tested and applied the same way every time; a wiki checklist drifts, gets skipped under deadline pressure, and leaves no machine-readable audit trail.
policy/ deny_root.rego deny_root_test.rego # CI: opa test policy/ && conftest test manifests/ -p policy/ # the same package can back a Gatekeeper ConstraintTemplate in the clusterLink to this question
All three, as layers. Conftest in CI for fast feedback, Gatekeeper or Kyverno at admission as the cluster backstop, and OPA or an authz service at runtime for request decisions. Same intent at multiple gates is defense in depth.
CI → conftest / opa test admission → Gatekeeper or Kyverno runtime → OPA sidecar / envoy external authzLink to this question
It's the stage after authn/authz where admission controllers can change or reject an object before it's persisted: built-in plugins, mutating and validating webhooks, and in-process CEL policies (ValidatingAdmissionPolicy, GA in 1.30; MutatingAdmissionPolicy, GA in 1.36). That's the hook point policy engines use to block non-compliant resources.
API request → authn → authz → mutating admission (webhooks, MutatingAdmissionPolicy)
→ object schema → validating admission (webhooks, ValidatingAdmissionPolicy) → etcdLink to this questionOPA is the general policy engine — Rego against any JSON-ish input. Gatekeeper wraps OPA for Kubernetes admission with ConstraintTemplates, Constraints, and audit. Kyverno is Kubernetes-native policy as YAML resources with CEL expressions (ValidatingPolicy, MutatingPolicy, GeneratingPolicy, ImageValidatingPolicy), and it's often faster to adopt if you're pure k8s.
OPA/Conftest → Terraform plans, Dockerfiles, multi-system Gatekeeper → Rego reuse + k8s admission + audit Kyverno → k8s-only teams wanting YAML + mutate/generate VAP (CEL) → simple per-object checks, no webhook to runLink to this question
It's OPA's declarative policy language. Rules query structured input and data and produce decisions — allow, deny messages, whatever you define. Unmatched rules contribute nothing; defaults give you the fallback.
package example default allow := false allow if input.method == "GET"Link to this question
You're writing a rule that adds a message to a deny or violation set when conditions on input hold; in Rego v1 that's deny contains msg if { ... }, since the old deny[msg] { } form no longer compiles. Multiple rules with the same name union their messages — so you get small independent checks instead of one giant if/else.
package k8s
deny contains msg if {
input.kind == "Pod"
c := input.spec.containers[_]
not c.securityContext.runAsNonRoot
msg := sprintf("container %v must set runAsNonRoot", [c.name])
}Link to this questionThere aren't explicit loops — [_] and variables iterate collections. Every expression in the rule body has to succeed, so it's logical AND. default fills in a value when nothing else matches.
# true if ANY container exposes port 22
ssh_exposed if {
input.spec.containers[_].ports[_].containerPort == 22
}Link to this questionMutating webhooks can change the object — inject sidecars, set defaults — before validation. Validating webhooks only accept or reject. I use both: mutate for paved roads, validate for the hard requirements.
mutating webhooks → defaults/sidecars validating webhooks → allow/deny final objectLink to this question
It's the built-in, in-process way to validate objects with CEL, stable since Kubernetes 1.30: a ValidatingAdmissionPolicy holds the rules and a binding sets scope and actions (Deny, Warn, Audit). There's no webhook to run, so no webhook outage or added latency. I use it for simple per-object checks and keep an engine for referential or external-data policies, audit of existing objects, and reporting.
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata: { name: replicas-limit }
spec:
policyName: replicas-limit
validationActions: [Warn, Audit] # later: [Deny]
# the ValidatingAdmissionPolicy holds matchConstraints + CEL validationsInterviewer often follows with: How can Gatekeeper's K8sNativeValidation engine let one ConstraintTemplate enforce through VAP while audit still runs in Gatekeeper?
Link to this questionConftest, OPA tests & data
It's for running Rego against structured config in CI — Kubernetes YAML, Dockerfiles, Terraform plan JSON — so violations fail the build before anything reaches a cluster.
conftest test deployment.yaml -p policy/ terraform show -json tf.plan | conftest test - -p policy/Link to this question
I'd use opa test with _test.rego files that assert deny or allow against sample inputs via with input as. Treat policies like code: tests in CI, coverage on critical rules, no silent enforcement gaps.
test_denies_root if {
count(deny) > 0 with input as {
"kind": "Pod",
"spec": {"containers": [{"name": "c"}]}
}
}
# opa test policy/ -vLink to this questioninput is the request or object under evaluation. data is external context — allowed registries, exemptions, org config — loaded into OPA. I keep allowlists in data so rules stay generic and updates are data PRs, not Rego rewrites.
# data.allowed_registries = ["registry.example.com/"]
deny contains msg if {
img := input.spec.containers[_].image
not strings.any_prefix_match(img, data.allowed_registries)
msg := sprintf("%v not from allowed registry", [img])
}
# not startswith(img, data.allowed_registries[_]) is rejected:
# a variable inside a negation must be bound outside itLink to this questionI'd freeze behavior with golden tests from real fixtures — pass and fail cases — add CI opa test plus coverage on deny paths, then refactor behind those tests. New rules ship with tests first; flaky policies stay in audit mode until they're proven.
Without tests, every policy PR is a production incident waiting to happen. I'd capture fixtures from actual denied AdmissionReviews and known-good manifests, and assert both that deny fires when it should and that compliant input yields an empty deny set. Table-driven cases help. For Gatekeeper, test the library Rego inside ConstraintTemplates the same way, and put a staging dryrun Constraint in place so you see violations before enforce. Document the data schemas too — broken data is as dangerous as broken Rego.
opa test policy/ -v --threshold 90 conftest test testdata/good/ -p policy/ # expect 0 conftest test testdata/bad/ -p policy/ # expect fails
Interviewer often follows with: How would you test policies that depend on cluster inventory like existing Namespaces or Images?
Link to this questionI'd ship versioned bundles — OCI or a bundle server — that OPA or Gatekeeper pulls, and export decision logs to a central sink. One source of truth for rules and data, plus an audit trail of allow and deny.
Bundles let you promote policy like any other artifact: build, sign, pull on agents. Pin versions per environment so a bad policy rolls back by pointing at the previous digest. Decision logs — carefully redacted — answer why something was denied and feed compliance evidence. Watch cardinality: logging every high-volume allow gets expensive, so I'd sample allows and keep all denies. Gatekeeper's audit scan of existing objects is how you see backlog, not just new requests.
opa run -s --set services.reg.url=https://reg.example --set services.reg.type=oci \ --set bundles.main.service=reg --set bundles.main.resource=reg.example/bundles/policy:1.4.2 # decision_logs → SIEM; alert on spike in denies
Interviewer often follows with: How would you sign and verify policy bundles so agents can't be fed malicious Rego?
Link to this questionI'd invest in clear deny messages, example fixtures in the repo, and a short CONTRIBUTING doc. Prefer many small rules with good msg strings over clever one-liners — the message is the UX.
msg := sprintf("%v/%v: set securityContext.runAsNonRoot=true (see docs/sec.md)", [input.kind, name])Link to this questionopa check catches parse and compile errors; opa fmt keeps style consistent. Together with opa test they're the minimum lint gate before merging policy changes.
opa check policy/ opa fmt --diff policy/ opa test policy/ -vLink to this question
Gatekeeper & Kyverno
A ConstraintTemplate defines the reusable Rego and registers a CRD. A Constraint instantiates that CRD with parameters and a match scope — kinds, namespaces, labels. Logic stays separate from where and how strictly it applies.
kind: K8sRequiredLabels
metadata: { name: ns-needs-owner }
spec:
match: { kinds: [{ apiGroups: [""], kinds: ["Namespace"] }] }
parameters: { labels: ["owner"] }Link to this questionI'd ship the Constraint in dryrun first, read the violation inventory, fix or time-box exemptions, scope enforce to a pilot namespace, then widen. I never land a new blocking policy cluster-wide in deny on day one.
dryrun and warn let admission report without rejecting. Audit periodically scans existing objects so you see debt, not just new creates. Fix workloads through GitOps, use data-driven exemptions with owners and expiry, then flip to deny for one team namespace. Watch initContainers, ephemeral containers, and controllers that recreate noncompliant pods. Pair with PSS labels where they overlap — don't double-fail people with conflicting messages.
kind: K8sPSPAllowedUsers
spec:
enforcementAction: dryrun # inventory first
parameters: { runAsUser: { rule: MustRunAsNonRoot } }
# kubectl get k8spspallowedusers -o yaml # status.violations
# later: enforcementAction: deny (scoped match.namespaces)Interviewer often follows with: How would you handle a controller that must create privileged system pods?
Link to this questionKyverno policies are Kubernetes resources with CEL expressions, one type per job: ValidatingPolicy blocks, audits, or warns; MutatingPolicy injects defaults or sidecars; GeneratingPolicy creates companion resources like a default NetworkPolicy per namespace; ImageValidatingPolicy checks signatures. The legacy ClusterPolicy with validate, mutate, generate, and verifyImages rules is deprecated in Kyverno 1.19. You don't need Rego for the common k8s guards.
Validate is the admission gate. Mutate can make the secure path automatic — inject runAsNonRoot, drop capabilities — so developers aren't fighting the platform. Generate keeps secondary resources in sync when a Namespace appears. Mutating policies need care around order, idempotency, and not fighting other mutators like mesh injectors. I prefer mutate for defaults and validate for hard requirements. Background scans catch existing debt the same way Gatekeeper audit does.
apiVersion: policies.kyverno.io/v1
kind: ValidatingPolicy
metadata: { name: require-non-root }
spec:
validationActions: [Deny]
matchConstraints:
resourceRules:
- apiGroups: [""]
apiVersions: [v1]
operations: [CREATE, UPDATE]
resources: [pods]
validations:
- message: "runAsNonRoot is required"
expression: >-
object.spec.containers.all(c,
c.?securityContext.?runAsNonRoot.orValue(false))Interviewer often follows with: When would you choose Kyverno generate over a Helm chart creating the same NetworkPolicy?
Link to this questionDry-run and audit report violations without blocking, which gives you a safe rollout and a visible backlog: Gatekeeper records them in the Constraint status, Kyverno in PolicyReports. Enforce rejects offending requests at admission. I ship audit, fix the workloads, then enforce, ideally namespace-scoped first.
# Gatekeeper Constraint spec: enforcementAction: dryrun # then warn, then deny # Kyverno ValidatingPolicy spec: validationActions: [Audit] # later [Deny] on the same policy # legacy ClusterPolicy used validationFailureAction: Audit | Enforce (deprecated)Link to this question
Policies verify image signatures and provenance — Kyverno ImageValidatingPolicy, Sigstore policy-controller, Gatekeeper plus cosign — and restrict registries so unsigned or untrusted artifacts never schedule.
CI verification is necessary but bypassable with direct kubectl or a rogue pipeline. Admission is the hard backstop: verify cosign signatures or keyless identities, optionally require SLSA provenance attestations, and deny :latest or unknown registries. Cache verification carefully for latency. Fail closed in prod, but keep a break-glass Namespace with loud audit. Combine with PSS restricted so even a signed image can't request privileged escalations.
kind: ImageValidatingPolicy # policies.kyverno.io/v1
spec:
validationActions: [Deny]
matchConstraints:
resourceRules: [{ apiGroups: [""], apiVersions: [v1], operations: [CREATE, UPDATE], resources: [pods] }]
matchImageReferences: [{ glob: "registry.example.com/*" }]
attestors:
- name: cosign
cosign: { key: { data: "-----BEGIN PUBLIC KEY-----..." } }
validations:
- expression: >-
images.containers.map(image, verifyImageSignatures(image, [attestors.cosign])).all(e, e > 0)
message: image signature verification failedInterviewer often follows with: Keyless vs key-based signing — what do you pin in the policy?
Link to this questionPSS is built-in admission — privileged, baseline, restricted — via namespace labels. It's a coarse floor. Gatekeeper and Kyverno add org-specific rules: labels, registries, mutates. I use PSS for the floor and policy engines for the rest, and I try not to duplicate the same check in three places.
kubectl label ns team-a \ pod-security.kubernetes.io/enforce=restricted \ pod-security.kubernetes.io/enforce-version=latestLink to this question
I'd use kyverno test with a local directory of policies, resource fixtures, and expected pass/fail results, and run it in CI on every policy PR — same discipline as opa test for Rego.
kyverno test ./policies # policies/... + kyverno-test.yaml with resources & resultsLink to this question
I'd group by Constraint and namespace, fix platform defaults first because they multiply, open team tickets with sample offending objects, add time-boxed exemptions only where needed, then re-audit before flipping enforce.
Dump Constraint status and violations, get them into a dashboard, and rank by what actually blocks enforce. Prefer mass fixes via GitOps base chart changes over one-off kubectl patches. Watch for false positives from system controllers. Communicate a freeze date for enforce. If most violations are one chart, patch that chart once. Keep dryrun on until the backlog is below an agreed threshold and on-call has a break-glass path.
kubectl get constraints kubectl get k8srequiredlabels -o json | jq '.items[].status.totalViolations'
Interviewer often follows with: How would you avoid audit noise from completed Jobs or Pods that no longer matter?
Link to this questionGovernance, exemptions & app authz
Exemptions are data in Git — namespaced allowlists or labels — PR-reviewed, owned, time-bounded, and logged when used. I never treat 'turn off the Constraint in the cluster' as the fix.
I'd model data.exemptions or Constraint match exclusions with an owner annotation and expiresOn. CI fails if expiry is missing or past. Decision logs or Kyverno reports should emit when an exemption matched. Security reviews the exemption backlog weekly; expired entries auto-deny. For break-glass, a short-lived ClusterRole and a temporary PolicyException with an expiry beats deleting the policy. Interviewers listen for exemptions-as-code and expiry — not 'we exclude kube-system and move on.'
# data/exemptions.yaml (committed)
namespaces:
- name: payment-legacy
owner: payments-oncall
reason: needs hostPath for HSM
expiresOn: "2027-03-31"Interviewer often follows with: How would you stop teams from rubber-stamping exemption PRs forever?
Link to this questionCI with Conftest gives fast developer feedback and cheap fails, but it can be skipped. Admission can't be bypassed for API writes and catches out-of-band applies. I run the same policy source at both gates.
Drift between CI Rego and Gatekeeper templates is a common failure — CI green, admission red, or CI so strict people learn to ignore it. Generate both from one library package, or share the same .rego via Conftest and ConstraintTemplate sync in GitOps. Watch the Rego version: Conftest and OPA 1.x parse v1 syntax by default, while Gatekeeper templates stay on v0 unless the template uses code with engine: Rego and source version v1. Measure bypasses: alert on kubectl applies from humans in prod namespaces. Preview environments should use the same Constraints as prod, dryrun or enforce as appropriate.
policy/lib/k8s.rego # shared # CI: conftest test -p policy/ # cluster: ConstraintTemplate embeds the same rules
Interviewer often follows with: What about mutations that only exist after admission, like injected sidecars — where do you validate those?
Link to this questionServices query OPA — sidecar, library, or a central PDP — with request context and get allow/deny plus any obligations. Authz policy and data update independently of app deploys, so fine-grained rules leave the codebase.
Patterns I've used: Envoy ext_authz calling OPA, an embedded SDK with policy bundles, or central OPA with careful latency SLOs. Push identity — JWT claims, SPIFFE ID — into input, keep role bindings in data, and version the bundles. Partial evaluation and caching matter at high RPS. Testing uses the same opa test fixtures as infra policy. Pitfalls: inconsistent input schemas across services, allow-by-default mistakes, and sync delay when data changes. Pair with decision logs so you can audit who accessed what.
package authz
default allow := false
allow if {
input.method == "GET"
input.user.roles[_] == "reader"
}Interviewer often follows with: How would you canary a new authz policy without locking everyone out?
Link to this questionI'd set PSS restricted as the floor, Gatekeeper or Kyverno for org rules — registries, labels, privileges — AppProjects and namespace budgets for tenancy, data-driven exemptions with expiry, dryrun-to-enforce rollouts, and shared Rego or YAML libraries tested in CI and synced via GitOps.
Tenancy needs layers: namespace-as-a-service, ResourceQuota and LimitRange, NetworkPolicy default-deny, PSS enforce=restricted, and custom Constraints for company red lines — no privileged, no hostPath, signed images only. Platform owns the cluster-wide policies; teams can get namespaced ones (Kyverno NamespacedValidatingPolicy, for example) for extra strictness, not weaker. I'd want violation dashboards per team, an SLO on admission latency, and break-glass runbooks. verifyImages on prod namespaces. Avoid one mega-Constraint — compose small, tested rules. Document the paved road so the secure path is the easy chart default.
1. PSS restricted (namespace labels) 2. Kyverno/Gatekeeper org Constraints (dryrun→deny) 3. verifyImages for prod 4. exemptions.yaml with owners + expiry 5. CI conftest same ruleset
Interviewer often follows with: How would you onboard a team that needs a privileged DaemonSet for a device plugin?
Link to this questionWebhook downtime can fail-open or fail-closed depending on config. Slow webhooks delay every API write. Overly broad match rules break system namespaces. Mutations can fight each other. And dryrun debt can hide until enforce day.
failurePolicy Ignore vs Fail is a conscious tradeoff (Gatekeeper ships with Ignore, so it fails open unless you change it): Fail protects compliance but can brick cluster ops if the webhook is down — so run HA replicas, anti-affinity, and monitor webhook latency and errors. Exclude kube-system carefully; don't blanket-exclude everything noisy. Watch for Constraints that deny the policy engine's own updates. Load-test admission under GitOps sync storms. Keep a break-glass admin path documented and tested. The four things to name are HA, failurePolicy, system exclusions, and sync-storm latency.
kubectl get validatingwebhookconfigurations kubectl -n gatekeeper-system get pods # alert: p99 of apiserver_admission_webhook_admission_duration_seconds > 1s # histogram_quantile(0.99, sum by (le, name) (rate(apiserver_admission_webhook_admission_duration_seconds_bucket[5m])))
Interviewer often follows with: Would you choose fail-open or fail-closed for a verifyImages webhook in prod?
Link to this questionI'd start with metrics and audit-only policies, ship secure defaults in golden charts, give teams self-service fixes, time-box exemptions, then enforce namespace-by-namespace with clear SLAs. Policy as a platform product — not a surprise club.
Change management beats Rego cleverness. Publish the roadmap, show violation dashboards per team, offer office hours, and fix the paved road so compliant deploys are easier than noncompliant ones. Enforce on new namespaces first, then backfill. Pair with PSS warn then enforce. Celebrate reduction in violations, not number of denies. Executive sponsorship matters when a deadline forces enforce. Empathy, metrics, and gradualism are the expert signal.
phase 1: Audit + dashboards phase 2: mutate defaults in chart templates phase 3: Enforce on greenfield ns phase 4: Enforce on brownfield after debt < threshold
Interviewer often follows with: What KPI proves the program is working besides number of Constraints?
Link to this questionNegation gets weird with undefined values and iteration — rules may not fire when fields are missing. I prefer positive checks, explicit default, and unit tests that cover the missing-field case.
In Rego, not p is true when p can't be proven. Combined with partial objects, 'deny if not runAsNonRoot' can miss containers where securityContext is entirely absent unless you structure helpers carefully. I use helpers that normalize to false when unset, or object.get with defaults. Always test missing, false, and true. It's a classic deep-dive for anyone claiming Rego fluency.
is_non_root(c) if c.securityContext.runAsNonRoot == true
deny contains msg if {
c := input.spec.containers[_]
not is_non_root(c)
msg := sprintf("%v must run as non-root", [c.name])
}Interviewer often follows with: How would you treat initContainers and ephemeralContainers in the same policy?
Link to this questionI'd use a break-glass admin path or carefully exempt the system namespace, fix the Constraint match exclusions, then add CI dry-run of Constraint changes against platform manifests before enforce.
Self-lockout happens when Constraints match control-plane or gatekeeper-system resources. Recovery is apiserver break-glass with local creds, remove or relax the Constraint, and be mindful of failurePolicy during the emergency. Prevention: excludedNamespaces for kube-system and gatekeeper-system — carefully — label-based exemptions for controllers, and policy tests that include the engine's own manifests. Never ship a new Constraint at enforce without dry-run metrics. Name the recovery path before you write the deny rule.
spec:
match:
excludedNamespaces: ["kube-system", "gatekeeper-system"]
# still monitor what you excludeInterviewer often follows with: How would you avoid over-excluding and creating a policy-free zone?
Link to this questionI'd parameterize Constraints per namespace or label — ConstraintTemplates plus per-tenant Constraints — or use Kyverno policies scoped by namespace selectors, with a default-deny baseline and explicit grants.
One global deny hostPath breaks legitimate device plugins for team A. Patterns that work: namespace labels selecting different Constraints, PolicyExceptions with expiry, separate Template parameters for allowedNamespaces. Default PSS restricted plus exemptions for node agents only. Governance question: who may create exceptions. Parameterization and tenancy labels beat copy-pasted Rego forks.
spec:
match:
namespaceSelector:
matchLabels: { tenant: "edge-devices" }
parameters:
allowHostPath: trueInterviewer often follows with: How would you prove team B never received the hostPath exemption?
Link to this questionI'd compare inputs. CI often tests raw YAML while admission sees mutated objects with defaults. Align the libraries, run Gatekeeper dry-run or test, and feed admission-shaped fixtures into Conftest.
Usual divergences: different Rego packages (Conftest only evaluates package main unless you pass --namespace or --all-namespaces, so CI can pass because nothing ran), missing data documents, Kubernetes defaulting like service account token mounts, and webhook mutation order. Fix it with shared policy modules as a versioned bundle, CI fixtures from kubectl apply --dry-run=server -o yaml, and Gatekeeper constraint status for violations. Version-pin the bundle in both places. Add a staging apply that must pass Gatekeeper before prod, and for critical charts optionally run a Kind cluster with Gatekeeper in PR CI. Document the input contract in the policy repo README: 'same policy' means same input schema and data documents, not just the same repo folder name.
kubectl apply --dry-run=server -o yaml -f deploy.yaml > admission-shaped.yaml conftest test admission-shaped.yaml -p policy/
Interviewer often follows with: How do you keep Conftest and Gatekeeper on the same bundle version in GitOps?
Link to this questionFor image provenance I prefer fail-closed with HA webhook replicas and a documented break-glass. Temporary fail-open is an explicit risk acceptance with monitoring — never a silent default.
failurePolicy=Fail bricks deploys when the webhook is down; Ignore lets unsigned images through. I'd design for 3+ replicas, a PDB, multi-AZ, SLOs on webhook latency, and a break-glass Namespace exemption owned by security. Cache or allowlist last-known-good digests if the product supports it. Pair with cluster PSA and network controls so one broken webhook isn't your only defense. State the threat model — supply-chain vs availability — and choose consciously.
kubectl -n cosign-system get deploy,pdb # break-glass: temporary Namespace label exempt=true with ticket + expiry
Interviewer often follows with: What secondary control still blocks :latest public images if verify-images is fail-open?
Link to this questionI usually suspect a too-broad deny that is always true, a wrong input path for the ConstraintTemplate, or a missing matcher. I'd roll back the Template, reproduce with gator or test fixtures, and require canary enforce on one Namespace first.
Templates that deny when input.review.object fields are undefined can fire on everything. Tests that only cover happy JSON miss the admission envelope shape — review.object.spec and friends. Process: deploy Template in dryrun, watch violation counts, enforce on a pilot Namespace, then fleet-wide. Keep the previous Template version for rollback. Admission envelope literacy plus progressive enforce is the expert signal.
# 1) dryrun Constraint cluster-wide # 2) enforce only ns/pilot # 3) compare allowed deploys before expanding
Interviewer often follows with: How would you structure Rego helpers to avoid deny-all when a field is missing?
Link to this questionReal-world scenarios
I'd treat it as an admission outage: check webhook endpoints and Gatekeeper pods, decide consciously between temporary failurePolicy=Ignore versus break-glass, restore webhook health, then re-enable fail-closed with an incident note.
When the API server can't get a timely ValidatingWebhookConfiguration response and failurePolicy is Fail, matching creates and updates get rejected — including platform work. First fifteen minutes: get the validatingwebhookconfiguration, endpoints, and Gatekeeper Deployment/Pod logs; check CPU, memory, and apiserver timeoutSeconds. Emergency paths: scale healthy replicas, fix network or DNS to the service, or briefly set failurePolicy=Ignore with explicit risk acceptance that noncompliant objects may land. Prefer restoring the webhook over living fail-open. After recovery: PDB, multi-replica, latency SLOs, and alert on webhook errors — not only on Constraint denies.
kubectl get validatingwebhookconfiguration -o wide kubectl -n gatekeeper-system get deploy,po,ep kubectl -n gatekeeper-system logs deploy/gatekeeper-controller-manager --tail=100 # last resort (ticketed): patch failurePolicy Ignore → restore Fail
Interviewer often follows with: What's the difference between a Constraint deny and a webhook timeout failure from the user's point of view?
Link to this questionI'd roll enforcementAction back to dryrun or warn immediately instead of deleting the Constraint, unblock deploys, inventory the real violations from audit, then re-enforce on a pilot namespace once owners are fixing debt.
Instant cluster-wide deny is a process failure, not a Rego talent show. Recovery: GitOps revert or kubectl patch enforcementAction to dryrun, communicate the freeze window, and use Gatekeeper audit status to list violators. Keeping the Constraint in dryrun keeps that violation list; deleting it loses the inventory along with the control. Re-enforce one canary Namespace at a time. The lasting fix belongs in the promotion pipeline, so enforce flips cannot skip dryrun soak, a canary Namespace, and CODEOWNERS review.
kubectl patch K8sRequiredLabels ns-needs-owner --type=merge \
-p '{"spec":{"enforcementAction":"dryrun"}}'
# then: kubectl get k8srequiredlabels ns-needs-owner -o jsonpath='{.status.totalViolations}'Interviewer often follows with: What CI gate would have caught this before the Constraint hit the cluster?
Link to this questionI'd inventory exceptions with owners and expiry, expire orphans, convert permanent skips into scoped Constraints or PSS labels, and require a ticket plus TTL for every new exemption.
Exemption sprawl is policy theater — the deny path looks strict while half the fleet is carved out. Method: export all PolicyExceptions, match.excludedNamespaces, and data.exemptions, join to CODEOWNERS or namespace labels, and mark anything without an owner or past expiry for deletion in waves. Prefer fixing the workload or parameterizing the Constraint over eternal carve-outs. Governance: PR template fields for risk, expiry, compensating control; a weekly stale-exception report; and deny creating exceptions without those fields — yes, with another policy.
# Kyverno PolicyException (policies.kyverno.io/v1, sketch)
metadata:
annotations:
owner: "payments-platform"
expires: "2027-03-31"
ticket: "SEC-4412"
spec:
policyRefs: [{ name: require-non-root, kind: ValidatingPolicy }]
matchConditions:
- name: legacy-namespace
expression: "object.metadata.namespace == 'payment-legacy'"Interviewer often follows with: How would you prove an exemption wasn't used as a silent permanent allow-all?
Link to this questionI'd profile which ConstraintTemplates are hot, simplify Rego — avoid nested comprehensions over huge inventories — move heavy checks to audit or CI, and raise capacity only after the policy is lean.
Admission is on the critical path: every create pays for every matching Constraint. Classic killers are iterating cluster inventory in Rego, unbounded comprehensions, and calling external data synchronously. Fix: refresh data documents out-of-band, push expensive inventory checks to Gatekeeper audit or Conftest in CI, split Templates so rarely-needed rules are narrowly matched, and watch webhook duration histograms. Horizontal scale only helps after the algorithmic waste is gone. Treat Rego like production code with budgets and fixtures that assert evaluation time.
spec:
match:
kinds: [{ apiGroups: ["apps"], kinds: ["Deployment"] }]
namespaces: ["payments"] # not cluster-wide if avoidable
# move "compare to all live Services" checks to audit/CIInterviewer often follows with: Why can a Constraint that is correct still be unsafe to run at admission?
Link to this questionI'd fix webhook reinvocation and ordering, make mutations idempotent and compatible with the mesh's expected pod shape, and move hard requirements to validate after all mutators finish.
MutatingAdmissionWebhook order and reinvocationPolicy decide whether Kyverno or Istio/Linkerd wins. Symptoms: lost labels, missing sidecars, or validate denying the post-mutate object. Resolve by documenting the intended end-state, setting reinvocation so validators see the final object, avoiding fields the mesh owns, and preferring mutate-for-defaults plus validate-for-invariants. Integration tests through the live webhook chain catch what unit tests miss. Admission is a pipeline — design the final object, not isolated policies.
# Kyverno: mutate defaults (runAsNonRoot) when unset # Istio: inject sidecar # Kyverno validate: deny if any container runs as root (final object) # reinvocationPolicy: IfNeeded on mutators
Interviewer often follows with: What breaks if you validate in a webhook that runs before the mesh injector?
Link to this questionI'd keep restricted for app namespaces, move privileged node agents to a dedicated namespace with baseline or privileged set intentionally, and never weaken restricted fleet-wide for one DaemonSet.
PSS restricted correctly rejects privileged, hostPath, and many initContainer patterns. Node agents — CSI, networking, tuners — often need privileged or host namespaces. That's a tenancy and placement problem, not a reason to drop restricted on payments. Pattern: kube-system or node-agent namespaces at privileged/baseline with tight RBAC and admission allowlists; app namespaces stay restricted. Document which controllers are exempt and why. Pair with Gatekeeper for finer rules inside the privileged NS.
kubectl label ns payments \ pod-security.kubernetes.io/enforce=restricted kubectl label ns node-agents \ pod-security.kubernetes.io/enforce=privileged # DaemonSet lives only in node-agents
Interviewer often follows with: Why is labeling the app namespace privileged to unblock one initContainer a bad trade?
Link to this questionI'd refuse a blind Friday flip: warn and audit first, fix or exempt with TTL, pilot enforce on low-risk namespaces, then wave the rest with a violation burn-down dashboard.
PSS modes — warn, audit, enforce — exist so you can see debt before you brick things. Export failing pods, prioritize easy wins like drop caps and non-root, isolate true privileged workloads, then enforce per Namespace wave. Communicate breakages early. Success metric is percent of namespaces at restricted enforce and exception age — not a calendar slogan. Change management plus the technical label strategy is what I want in an expert answer.
# week 1: warn+audit=restricted on all app ns # week 2: enforce on pilot ns # week 3+: enforce waves; exceptions expire ≤30d
Interviewer often follows with: What signal tells you a namespace is ready to move from audit to enforce?
Link to this questionI'd version policies as signed bundles or PRs, require dry-run apply plus violation budgets in CI and staging, gate enforce on metrics and human approval, and auto-revert on deny spikes.
Promotion path: PR → opa test/conftest → deploy Constraint at dryrun to staging → soak with zero unexpected denies → canary enforce Namespace → fleet. CI tests run against admission-shaped fixtures, and the soak is a set number of days in dryrun with a violation budget near zero. Enforce flips need a second reviewer from platform and security, via CODEOWNERS on the Constraint files. Watch webhook latency, deny rate, and audit violation count with burn alerts, and track time-in-dryrun as a platform metric. Rollback is pointing GitOps at the previous Constraint revision within minutes. Policy delivery is a CD problem with stricter gates than app code.
# CI must prove: # - opa test / kyverno test green # - dryrun violations ≤ budget OR owned exceptions # - CODEOWNERS approve enforcementAction: deny
Interviewer often follows with: How would you stop someone from kubectl-editing enforcementAction around GitOps?
Link to this questionI'd offer a time-boxed, monitored break-glass with logging, network isolation, and forced expiry — not a permanent policy-free zone — and escalate lasting exceptions to a risk register.
Permanent no-policy namespaces become the production path under pressure. Compromises I'll offer: short TTL Namespace with PSA privileged and Gatekeeper exempt, but NetworkPolicy default-deny egress, no internet-facing Services, mandatory runtime detection, and automatic teardown. Every use pages security. If leadership insists on permanence, residual risk and compensating controls go in writing. Negotiate safer velocity — don't rubber-stamp shadow IT.
ns/break-glass-YYYYMMDD psa: privileged gatekeeper: exempt (TTL annotation) netpol: default deny egress except approved CIDRs owner + expires + auto-delete CronJob
Interviewer often follows with: What compensating control is insufficient if break-glass pods can still reach the cloud metadata service?
Link to this questionI'd make deny messages actionable — what, why, doc link — make sure Gatekeeper aggregates causes clearly, and add a self-service 'why was this denied?' runbook tied to Constraint names.
Admission UX is part of policy adoption. Messages should name the Constraint, the offending field, and the fix. Avoid overlapping Templates that emit opaque Rego errors. Platform docs: kubectl describe plus Gatekeeper Constraint status. Consider warn mode during onboarding so engineers learn the messages before hard deny. I'd measure time-to-understand-deny in onboarding surveys.
msg := sprintf("%v: container %v must set runAsNonRoot=true (Constraint K8sPSPNonRoot; docs/runAsNonRoot.md)", [input.review.object.metadata.name, c.name])Interviewer often follows with: How would you test that deny messages stay stable across Rego refactors?
Link to this questionI'd treat generated resources as owned by the policy engine or fold the NetworkPolicy into the team's GitOps desired state — one owner — and alert on drift instead of silent regenerate wars.
Generate vs GitOps ownership conflicts cause thrash: Kyverno recreates, Argo or Flux deletes. Fixes: teams own the NetworkPolicy in git with a Conftest or Kyverno validate that it exists and matches baseline; or Kyverno generate with clear labels and Flux ignore rules — never both fighting. I often prefer validate 'must have NP matching baseline' so GitOps stays source of truth. Pick a single control plane for each object kind.
# team repo must include networking/allow-dns.yaml # CI: kyverno apply / conftest test # cluster: validate policy denies Namespace without labeled NP
Interviewer often follows with: When is generate still the right tool despite GitOps?
Link to this questionRelated
- Cheat sheetOPA & Rego cheat sheet
- CourseOPA & Rego
- CoursePolicy-as-code at scale
- CourseCheckov & IaC scanning
- Field noteOPA Gatekeeper: writing constraint templates from scratch
- Field notePolicy-as-code for Terraform: Trivy, Checkov and OPA in review
- Field notePod Security Standards: enforce restricted without breakage
Primary references
Found a technical issue on this page? Report it with the tool version you used and the behavior you saw. How resources are maintained.