Kubernetes RBAC least privilege: from admin to scoped roles
Replace blanket cluster-admin with namespaced Roles, audit who can do what, and test access with kubectl auth can-i.
kubectl get clusterrolebindings -o json | jq -r '.items[] | select(.roleRef.name=="cluster-admin") | .metadata.name + " " + ([.subjects[]? | .kind + ":" + (.namespace // "") + "/" + .name] | join(", "))'cluster-admin Group:/system:mastersjenkins-deployer ServiceAccount:ci/jenkinsdashboard-admin ServiceAccount:kube-system/dashboardplatform-oncall Group:/platform-teamfour subjects, each a full-cluster takeover if its credential leaks; only the first is expectedClusters accumulate cluster-admin bindings the way a garage accumulates boxes: a CI account that needed to create a namespace once, a dashboard from a tutorial, an on-call group that was going to be narrowed later. Least privilege does not mean zero cluster-admin bindings, since someone needs break-glass access; it means every binding is intentional, scoped to what its subject does, and provable with kubectl auth can-i. Getting there is a loop, and the first pass through it is the audit above.
The token proves which service account a request comes from; RBAC decides what that account may do. A leaked token is exactly as dangerous as the bindings behind it, so the bindings are what this audit tightens.
See the cluster through the subject’s eyes
kubectl auth can-i --list with --as prints every verb and resource a subject is allowed, resolved through all of its bindings. It is the view to take before writing any Role, because it shows the gap between what the subject is granted and what its workload actually calls. A in the verbs column is a full grant on that resource; a on both columns is cluster-admin under another name. Every authenticated subject also carries a handful of discovery rules (selfsubjectaccessreviews, the /api and /healthz non-resource URLs), which is why a service account with no bindings at all still lists something. kubectl auth whoami shows what the current credentials resolve to, which matters when the surprise is which identity a pipeline is using.
kubectl auth can-i --list --as=system:serviceaccount:ci:jenkinsResources Non-Resource URLs Resource Names Verbs*.* [] [] [*]everything, everywhere: the jenkins-deployer bindingkubectl auth can-i --list --as=system:serviceaccount:shop:web -n shopconfigmaps [] [] [get]pods [] [] [get list watch]a reader, in one namespace: what the workload needs and nothing elseNamespaced Role, bound to one subject
A Role exists in one namespace; a ClusterRole applies everywhere it is bound. Default to a Role unless the subject needs cluster-scoped resources such as Nodes, PersistentVolumes or CustomResourceDefinitions. The verbs come from what the workload does, not from what might be convenient: a deployer needs create, patch and get on Deployments in its namespace, and does not need delete on Secrets to do that. Write the verbs out; a wildcard is a decision not to decide.
apiVersion: rbac.authorization.k8s.io/v1kind: Rolemetadata:namespace: shopname: deployerrules:- apiGroups: ["apps"]resources: ["deployments"]verbs: ["get", "list", "patch", "create"]- apiGroups: [""]resources: ["configmaps"]verbs: ["get", "list", "create", "patch"]- apiGroups: [""]resources: ["pods"]verbs: ["get", "list"] # to watch a rollout; not delete, not exec---apiVersion: rbac.authorization.k8s.io/v1kind: RoleBindingmetadata:namespace: shopname: jenkins-deploys-shopsubjects:- kind: ServiceAccountname: jenkinsnamespace: ciroleRef:kind: Rolename: deployerapiGroup: rbac.authorization.k8s.io
A subject that deploys to five namespaces gets five RoleBindings to the same Role definition (or one ClusterRole bound into each namespace with a RoleBinding, which keeps the rules in one place while the grant stays namespaced). What it does not get is a ClusterRoleBinding, because that would make the same verbs valid in kube-system.
The verbs that are more than they look
Grants that quietly equal more than the resource they name
| Grant | Why it escalates | Give it to |
|---|---|---|
create on pods (or pods/exec) | a pod can mount any Secret in the namespace and run as any service account there | controllers that must create pods; never a human role that only reads |
get on secrets | every token and credential in the namespace, including other service accounts’ secrets | the specific workload, on named resources where possible (resourceNames) |
bind, escalate on roles | lets the subject grant itself permissions it does not hold | nobody, outside break-glass |
impersonate on users, groups, serviceaccounts | become anyone the API server knows | the identity gateway, if you run one |
edit or admin ClusterRoles bound to automation | aggregated roles grow as CRDs are added; automation ends up with verbs nobody reviewed | human developers in their own namespace; controllers get an explicit Role |
The built-in view, edit and admin ClusterRoles are aggregated: any ClusterRole labelled for aggregation, including ones an operator installs, is folded into them. That makes them reasonable for people working in a namespace and wrong for a controller, whose permission set should not change because someone installed a CRD.
Tokens that are never used should not exist
Every pod gets its service account's token mounted at /var/run/secrets/kubernetes.io/serviceaccount/token unless told otherwise, and most application pods never call the API. A compromised pod with a mounted token and a permissive binding is the standard path from RCE to cluster. automountServiceAccountToken: false on the ServiceAccount (or the Pod) removes the token from pods that do not need it; the ones that do get bound, time-limited tokens by default on current versions, which expire and are tied to the pod that received them.
apiVersion: v1kind: ServiceAccountmetadata:name: webnamespace: shopautomountServiceAccountToken: false # the web pods never call the API server
Prove the negative, then keep proving it
The check that matters is the one that should fail. After tightening, ask can-i for the dangerous verbs, not the granted ones, and put the same questions into CI so a later binding cannot quietly widen the account. The audit log is the other half: a create on a ClusterRoleBinding to cluster-admin is the event that undoes this work, and it should page someone.
kubectl get clusterrolebindings -o json | jq -r '.items[] | select(.roleRef.name=="cluster-admin") | .metadata.name + " " + ([.subjects[]? | .kind + ":" + (.namespace // "") + "/" + .name] | join(", "))'cluster-admin Group:/system:masterskubeadm:cluster-admins Group:/kubeadm:cluster-adminsa fresh kubeadm-built cluster: two groups, both expected. Anything else in this list is what the audit is forkubectl auth can-i --list --as=system:serviceaccount:ci:jenkins -n shop | grep -vE "selfsubject|well-known|/api|/healthz|/livez|/readyz|/openapi|/version|Non-Resource"configmaps [] [] [get list create patch]deployments.apps [] [] [get list patch create]pods [] [] [get list]for v in "delete pods" "get secrets" "create clusterrolebindings" "create pods --subresource=exec"; do printf "%-32s " "$v"; kubectl auth can-i $v --as=system:serviceaccount:ci:jenkins -n shop; donedelete pods noget secrets nocreate clusterrolebindings Warning: resource 'clusterrolebindings' is not namespace scoped in group 'rbac.authorization.k8s.io' nocreate pods --subresource=exec nokubectl auth can-i patch deployments --as=system:serviceaccount:ci:jenkins -n shopyesfour expected noes and one expected yes: the Role does what the workload needs and nothing an attacker would want. can-i exits 1 on a no, so the loop can fail a pipelineTOKEN=$(kubectl -n ci create token jenkins --duration=10m) && kubectl config set-credentials jenkins --token="$TOKEN" && kubectl config set-context jenkins --cluster=kind-secopslog-p2c-rbac --user=jenkinskubectl --context=jenkins auth whoami | grep UsernameUsername system:serviceaccount:ci:jenkinskubectl --context=jenkins -n shop get pods -o namepod/webkubectl --context=jenkins -n shop get secretsError from server (Forbidden): secrets is forbidden: User "system:serviceaccount:ci:jenkins" cannot list resource "secrets" in API group "" in the namespace "shop"kubectl --context=jenkins -n shop exec web -- idError from server (Forbidden): pods "web" is forbidden: User "system:serviceaccount:ci:jenkins" cannot create resource "pods/exec" in API group "" in the namespace "shop"kubectl --context=jenkins get nodesError from server (Forbidden): nodes is forbidden: User "system:serviceaccount:ci:jenkins" cannot list resource "nodes" in API group "" at the cluster scopethe same answers from the authorizer itself, with a real token in a context that carries nothing else: delete, another namespace and the cluster scope were refused the same way (exit 1 each). get pods does not include exec: pods/exec is its own resourcekubectl create clusterrolebinding jenkins-deployer --clusterrole=cluster-admin --serviceaccount=ci:jenkinskubectl auth can-i get secrets --as=system:serviceaccount:ci:jenkins -n shop; kubectl auth can-i delete pods --as=system:serviceaccount:ci:jenkins -n kube-systemyesyeskubectl get clusterrolebindings -o json | jq -r '…' | grep jenkins # the audit from the top of the pagejenkins-deployer ServiceAccount:ci/jenkinskubectl delete clusterrolebinding jenkins-deployerclusterrolebinding.rbac.authorization.k8s.io "jenkins-deployer" deletedkubectl --context=jenkins -n shop get pods -o name >/dev/null && kubectl --context=jenkins -n shop get secrets 2>&1 | grep -o "secrets is forbidden"secrets is forbiddenone object to delete, and the same token immediately answers the way the Role says: the reversal is verified with the request that should fail, not the one that should workWhen a binding is wrong: recovery in order of preference
| Symptom | Diagnosis | Reversal and verification |
|---|---|---|
| the cluster-admin audit shows a subject nobody expected | kubectl get clusterrolebinding <name> -o yaml for who created it and when (the audit log has the request) | kubectl delete clusterrolebinding <name>; verify with auth can-i get secrets --as=<subject> -n <ns> returning no and, if the subject is a service account, a real request with its token (executed above) |
| a pipeline stopped working after a Role was tightened | kubectl auth can-i --list --as=<subject> -n <ns> shows the verb it now lacks; the job log has the 403 with the exact resource | add the one verb the workload calls to the Role (never a wildcard), kubectl apply, re-run the job; the negative loop must still print four noes (documentation-backed) |
| a token from a mounted secret leaked | the binding behind it is the blast radius: auth can-i --list as that service account | delete the binding first, then rotate: kubectl delete secret <token-secret> for legacy token secrets, or a new bound token is minted on the next Pod start; set automountServiceAccountToken: false where the pod never calls the API (the unmounted case was observed; rotation is documentation-backed) |
RBAC decides what a subject may ask the API server. It does not decide what a pod may connect to (NetworkPolicy) or what it may do on the node (Pod Security). The three are usually tightened in that order, and the audit loop here is the one that gets repeated most often, because bindings are what people add under pressure.