Pod Security Standards: enforce restricted without breakage
Roll out the restricted profile namespace by namespace, find violators with audit mode first, and carve exceptions that expire.
kubectl apply -f deploy.yaml; echo exit=$?deployment.apps/api unchangedexit=0the Deployment is accepted: enforce applies to Pods, not to the templatekubectl get rs -n payments -l app=api -o jsonpath="{range .items[*]}{.metadata.name} replicas={.status.replicas} {.status.conditions[0].type}: {.status.conditions[0].message}{'\n'}{end}"api-7f987c4746 replicas=0 ReplicaFailure: pods "api-7f987c4746-mt2rj" is forbidden: violates PodSecurity "restricted:v1.35": allowPrivilegeEscalation != false (container "api" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "api" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "api" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "api" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")zero replicas come up, and the reason is on the ReplicaSet as a ReplicaFailure condition, not on the apply. The message names every field to changeThat failure mode is the whole argument for a phased rollout. Pod Security Admission enforces the three Pod Security Standards with one namespace label, and enforce rejects Pods, not the Deployment that creates them: kubectl apply succeeds, the ReplicaSet fails to create replicas, and the person who finds out is whoever notices the service has no endpoints. The warn and audit modes exist to move that discovery earlier, to the developer's terminal and to the audit log, while nothing is blocked.
Three modes per namespace, each set to one of three levels. The rollout below moves a namespace from warn+audit at restricted to enforce at restricted without a step where workloads are silently blocked.
Three profiles, and which namespaces get which
Pod Security Standards
| Level | Blocks | Typical namespaces |
|---|---|---|
privileged | nothing | CNI, storage drivers, node agents that genuinely need host access |
baseline | host namespaces, privileged containers, most hostPath volumes, added capabilities beyond a default set, host ports | ingress controllers and monitoring agents that need something restricted forbids, with a written reason |
restricted | everything baseline blocks, plus: must set runAsNonRoot, drop ALL capabilities (NET_BIND_SERVICE may be added back), allowPrivilegeEscalation false, a seccomp profile, only the safe volume types | application namespaces |
Step one: warn and audit at the target level
Label the namespace with warn and audit at restricted and change nothing else. warn returns a warning on every kubectl apply that would fail under enforcement, including applies of Deployments and other workload objects, because warn and audit evaluate the Pod template inside them. audit records the same violations as annotations on the API server's audit events, which is the complete list, including the workloads nobody has redeployed since the label went on. Both modes admit the object; in the run the non-compliant Deployment was created with the warning below and its replica was available a second later.
kubectl label ns payments pod-security.kubernetes.io/warn=restricted pod-security.kubernetes.io/warn-version=v1.35 pod-security.kubernetes.io/audit=restricted pod-security.kubernetes.io/audit-version=v1.35namespace/payments labeledkubectl apply -f deploy.yamlWarning: would violate PodSecurity "restricted:v1.35": allowPrivilegeEscalation != false (container "api" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "api" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "api" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "api" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")deployment.apps/api createdkubectl apply -f deploy-fixed.yaml # runAsNonRoot, RuntimeDefault seccomp, no privilege escalation, drop ALLdeployment.apps/api configuredno warning the second time: the four fields are the whole difference between the two filesapiVersion: v1kind: Namespacemetadata:name: paymentslabels:pod-security.kubernetes.io/warn: restrictedpod-security.kubernetes.io/warn-version: v1.37pod-security.kubernetes.io/audit: restrictedpod-security.kubernetes.io/audit-version: v1.37
Harvesting the list needs the audit log to be on; the violation text is in the pod-security.kubernetes.io/audit-violations annotation of each event. One sprint of normal deploys is usually enough for every workload in the namespace to have been evaluated at least once.
jq -r 'select(.annotations["pod-security.kubernetes.io/audit-violations"]) | .objectRef.namespace + "/" + .objectRef.resource + "/" + .objectRef.name + ": " + .annotations["pod-security.kubernetes.io/audit-violations"]' audit.log | sort -upayments/deployments/api: would violate PodSecurity "restricted:v1.37": allowPrivilegeEscalation != false (container "api"), runAsNonRoot != true (pod or container "api"), seccompProfile (pod or container "api") must be setpayments/cronjobs/reconcile: would violate PodSecurity "restricted:v1.37": unrestricted capabilities (container "job" must set securityContext.capabilities.drop=["ALL"])each line is one workload and the exact fields to changeStep two: fix the securityContext, not the policy
Every violation maps to a field. The pod-level securityContext takes runAsNonRoot: true and seccompProfile: { type: RuntimeDefault }; each container takes allowPrivilegeEscalation: false and capabilities: { drop: ["ALL"] }. An image that runs as root needs a numeric runAsUser in the manifest and a filesystem that user can use, which is an image change rather than a YAML change. For Helm-managed workloads the fix goes into the chart values or an upstream issue, otherwise the next helm upgrade reverts it.
spec:template:spec:securityContext:runAsNonRoot: truerunAsUser: 10001seccompProfile:type: RuntimeDefaultcontainers:- name: apisecurityContext:allowPrivilegeEscalation: falsecapabilities:drop: ["ALL"]
Step three: enforce, pinned to a version
When the audit log has been quiet for the namespace, add the enforce label. Pin enforce-version to a minor version rather than latest: the standards can tighten with a Kubernetes release, and latest means a cluster upgrade can start rejecting Pods that passed the day before. Keep warn and audit on latest so the upcoming tightening shows up as warnings first. Adding or changing an enforce label triggers a dry-run evaluation of the namespace's existing Pods, and the resulting warnings are your last chance to catch something before a Deployment rollout tries to replace those Pods. Two things about the version label from the run, on a 1.35.0 server: a label naming a version that server did not have yet (v1.37) was accepted without any warning and echoed in the violation text as restricted:v1.37, so in that run a wrong minor number was not caught by the API server; and 1.35 without the v was rejected outright with must be "latest" or "v1.x". The Pod Security Admission documentation says how a version is resolved; what the run shows is that the server will not catch the mistake for you, so pin deliberately to the version you actually run and review the label when the cluster is upgraded.
labels:pod-security.kubernetes.io/enforce: restrictedpod-security.kubernetes.io/enforce-version: v1.37pod-security.kubernetes.io/warn: restrictedpod-security.kubernetes.io/warn-version: latestpod-security.kubernetes.io/audit: restrictedpod-security.kubernetes.io/audit-version: latest
kubectl label ns payments pod-security.kubernetes.io/enforce=restricted pod-security.kubernetes.io/enforce-version=v1.35Warning: existing pods in namespace "payments" violate the new PodSecurity enforce level "restricted:v1.35"Warning: api-7f987c4746-d8h22 (and 1 other pod): allowPrivilegeEscalation != false, unrestricted capabilities, runAsNonRoot != true, seccompProfilenamespace/payments labeledthe dry run names the Pods that would not be admitted again; they keep running (Running in kubectl get pod) until something replaces themkubectl apply -f privileged-pod.yaml; echo exit=$?Error from server (Forbidden): error when creating "privileged-pod.yaml": pods "priv" is forbidden: violates PodSecurity "restricted:v1.35": privileged (container "priv" must not set securityContext.privileged=true), allowPrivilegeEscalation != false (container "priv" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "priv" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "priv" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "priv" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")exit=1apply_fixed_wait # fixture helper: kubectl apply -f deploy-fixed.yaml, then poll .status.availableReplicasavailable replicas: 1 after 2sthe fix: the ReplicaSet for the compliant template creates its Pod at onceunlabel_apply_bad_wait # fixture helper: remove the two enforce labels, apply deploy.yaml, count its warning lines, poll availableReplicas1available replicas: 1 after 0sthe other way out, for a namespace that must ship now: remove the enforce labels and the non-compliant template is admitted again, still with the warning. It is a rollback of the control, not a fixWhen enforce bites: recovery in order of preference
| Symptom | Diagnosis | Reversal and verification |
|---|---|---|
a Deployment has zero available replicas after a rollout, kubectl apply said nothing | kubectl get rs -n <ns> -o jsonpath shows a ReplicaFailure condition with violates PodSecurity and the fields | apply the securityContext fields the message names; verify availableReplicas (executed: 1 within two seconds of the fixed apply) |
| the fix cannot ship today and the service is down | the namespace was moved to enforce with unfixed workloads | remove pod-security.kubernetes.io/enforce and enforce-version from the namespace; the ReplicaSet recreates the Pod immediately, warn and audit keep reporting (executed); put enforce back with the fix in the same change |
Pods keep running but the label change printed existing pods … violate | the dry run at label time found workloads not yet redeployed | nothing to revert yet: fix those workloads before their next rollout, node drain or crash replaces them (the dry-run warning was observed; the later replacement failure is the first row) |
| a namespace that must run privileged workloads is now blocked | the wrong level for that namespace | label it baseline or privileged with a written reason, or exempt the controller in the AdmissionConfiguration (documentation-backed; not executed) |
The right-hand column is where Gatekeeper or Kyverno come in, after PSA has set the floor. Their rollout discipline is the same one described above, under a different name: dry run, fix, then deny.