Turn on the Kubernetes audit log and actually read it

Write a Kubernetes audit policy that captures the right events at the right level, then query the log to find exactly who deleted that deployment at 2am.

Sep 24, 2024·Updated ·7 min readAdvanced·By SecOpsLog · documentation-verified

kubectl get events will not tell you who deleted a Deployment last night. Events are best-effort, expire after an hour by default, and describe what controllers did rather than who asked. etcd only holds the current state. The record you want is the API server's audit log: one JSON line per request stage with the authenticated user, the verb, the object reference, the source IPs and the response code. It is off by default because the naive configuration logs every request body of every watch, which is a storage problem before it is a security tool.

Turning it on well is mostly a matter of writing the policy so that it records identities and verbs everywhere, bodies only where a body is evidence, and nothing at all for the handful of callers that generate most of the volume. The policy is evaluated top to bottom and the first matching rule wins, so order is part of the design.

From request to audit record

Each stage of a request can produce an event; the policy decides its level; the backend decides where it lands. Most of the noise reduction happens in the middle column, most of the durability question in the right one.

Kubernetes audit pipeline: each request stage produces an event, the first matching policy rule sets the level from None to RequestResponse, and the event goes to the log file or webhook backend API requestkubectl, controlleror service accountkube-apiserverone event per stagePolicy: first matchrules read in orderStages (omitStages)1RequestReceived2ResponseStarted3ResponseComplete4PanicLevel from the ruleNonedroppedMetadatawho/verb/objectRequest+ request bodyReqResp+ response bodyBackendlogJSON lines file--audit-log-pathhookbatched POST--audit-webhook-*A rule below a matching one is never reached: level None noise drops go first, the catch-all last.ResponseStarted exists only for long-running requests such as watch.

A policy that records identity everywhere and bodies rarely

Four levels exist: None drops the event, Metadata records who, what, when and the outcome, Request adds the request body, and RequestResponse adds the response body too. Secrets belong at Metadata: the audit trail should say that system:serviceaccount:ci:deployer read shop/db-credentials at 02:14, not copy the credential into a second store with different access controls. RBAC changes are the case for Request, because the submitted Role or ClusterRoleBinding is the evidence. Everything else that mutates state gets Metadata. omitStages: [RequestReceived] at the top removes the duplicate event every request would otherwise emit before it is handled.

audit-policy.yaml
apiVersion: audit.k8s.io/v1
kind: Policy
omitStages:
- RequestReceived
rules:
# First match wins: drop known high-volume noise before any broad rule.
- level: None
users: ["system:kube-proxy"]
verbs: ["watch"]
- level: None
resources:
- group: ""
resources: ["events"]
# Record who touched a Secret, never the Secret itself.
- level: Metadata
resources:
- group: ""
resources: ["secrets"]
# For RBAC mutations the submitted object is the evidence; skip the response body.
- level: Request
verbs: ["create", "update", "patch", "delete"]
resources:
- group: "rbac.authorization.k8s.io"
resources: ["roles", "rolebindings", "clusterroles", "clusterrolebindings"]
# Every other mutation: identity, verb, object, outcome.
- level: Metadata
verbs: ["create", "update", "patch", "delete"]

Reads other than Secrets fall through every rule in this file, and a request that matches no rule is not logged. That is a deliberate starting point, not an oversight: adding - level: Metadata as a final catch-all is a one-line change once you know the volume the cluster produces, and it is much easier to widen a policy than to explain to storage why the audit log grew by a terabyte in a week.

RequestResponse on Secrets copies them into the audit store
Any rule that puts Secrets or ConfigMaps at Request or RequestResponse writes their contents into the log, the shipper, the SIEM and every backup of those. If one non-secret workflow really needs body-level evidence, scope Request to that resource and verb rather than raising the level broadly.
Go deeper in a courseKubernetes security (CKS-aligned)Audit policy is one lesson of the track; RBAC, Pod Security and runtime guards are the rest.View course

Two backends, one rule about the control plane disk

Self-managed clusters pass the policy with --audit-policy-file; managed clusters expose the same policy levels through a control-plane setting instead of flags. The log backend writes JSON lines to --audit-log-path, rotated by --audit-log-maxsize, --audit-log-maxbackup and --audit-log-maxage. The webhook backend posts batches to an HTTP endpoint described by a kubeconfig-shaped file in --audit-webhook-config-file. The reason to prefer the webhook, or a shipper that tails the file aggressively, is that audit volume spikes during exactly the incidents you want the log for, and a control plane node whose disk fills up stops serving the API.

kube-apiserver flags
--audit-policy-file=/etc/kubernetes/audit-policy.yaml
--audit-log-path=/var/log/kubernetes/audit/audit.log
--audit-log-maxsize=100 # MB before rotation
--audit-log-maxbackup=10
--audit-log-maxage=30 # days
# or, instead of (or in addition to) the file:
--audit-webhook-config-file=/etc/kubernetes/audit-webhook.kubeconfig

When the API server runs as a static Pod, the policy file and the log directory are on the host, so both need a hostPath volume mounted into the Pod. Forgetting the log directory mount is the usual reason a freshly configured backend writes nothing.

Reading it: who deleted the Deployment

A completed delete is an event with stage: ResponseComplete, verb: delete and the object reference, and user.username is the identity the authentication layer established, which for CI is a service account and for humans is whatever your OIDC provider put in the token. sourceIPs is the value to correlate with VPN or bastion logs when the username is not enough.

bash — who deleted the deployment
jq -c 'select(.verb=="delete" and .objectRef.resource=="deployments" and .objectRef.namespace=="shop")' audit.log
{"user":{"username":"system:serviceaccount:ci:deployer"},"objectRef":{"name":"api","namespace":"shop"},"stage":"ResponseComplete","sourceIPs":["10.0.4.22"],"responseStatus":{"code":200}}
jq -c 'select(.objectRef.resource=="secrets" and .verb=="get") | [.user.username, .objectRef.namespace, .objectRef.name, .requestReceivedTimestamp]' audit.log | tail -3
["system:serviceaccount:shop:api","shop","db-credentials","2026-09-12T02:14:09.311Z"]
Metadata level: the caller and the object, nothing from the Secret

The four patterns worth an alert

Once the log leaves the node, four queries cover most of the value: a clusterrolebindings create or patch that grants cluster-admin; any successful request from system:anonymous; Secret list or get calls from a service account that does not normally read Secrets in that namespace; and a burst of deletes in one namespace inside a short window. Separate the humans from the controllers when tuning: ci:deployer deleting Deployments all day is the job, whereas a human deleting a ClusterRole at 03:00 is a page.

What the audit log answers, and what it cannot
It answers
Who called the API, from where
Which object, which verb, which result
What RBAC object was submitted
When, to the millisecond
It does not answer
What a Pod did after it started
Kubelet-local actions (exec into the node)
Anything before logging was enabled
Why: that is still a conversation

The log is evidence after the fact. Stopping the same action next time is RBAC and admission work: least-privilege Roles for CI service accounts, a Gatekeeper or ValidatingAdmissionPolicy rule that blocks cluster-admin bindings outside a change window, and periodic kubectl auth can-i --list reviews. The Kubernetes security (CKS) track treats audit logging, RBAC and admission as one system, the way they behave in an incident.

Related posts

Quick reference