SELinux on OpenShift: stop chmod 777-ing your way out
Read AVC denials, write a targeted policy module, and keep SCCs strict instead of granting privileged.
ausearch -m AVC -ts recenttype=AVC msg=audit(1757667721.412:9314): avc: denied { read } for pid=2114 comm="nginx" name="index.html" dev="dm-0" ino=8391 scontext=system_u:system_r:httpd_t:s0 tcontext=unconfined_u:object_r:user_home_t:s0 tclass=file permissive=0ausearch -m AVC -ts recent | audit2whyWas caused by: Missing type enforcement (TE) allow rule. You can use audit2allow to generate a loadable module to allow this access.audit2why offers a module; the denial itself says the file is labelled as someone’s home directory, which is the actual problemA file with mode 0644 that a service cannot read is the moment most people type setenforce 0, and the problem goes away because the seatbelt is now unbuckled for every process on the host. The denial above is readable in one pass. The subject is nginx, running in the httpd_t domain; it was denied read on a file whose type is user_home_t. SELinux is not confused: a web server has no business reading a file typed as home-directory content. Somebody copied the site from /home/alice/www and the label came with it. The fix is the label.
Fix the label, persistently
chcon changes a label on disk and the change is lost at the next relabel, so it is the wrong tool for anything but an experiment. semanage fcontext records a rule in the policy ("everything under /srv/www is httpd_sys_content_t"), and restorecon applies the rules to the files that exist now. New files created under that path by any process get the right label from then on, and a full relabel produces the same result rather than undoing it.
# 1. record the rule: this path tree is web contentsemanage fcontext -a -t httpd_sys_content_t "/srv/www(/.*)?"# 2. apply it to what exists todayrestorecon -Rv /srv/www# 3. verify the type the service will seels -Z /srv/www/index.html# system_u:object_r:httpd_sys_content_t:s0 /srv/www/index.html
When the denial is not a label but a behaviour the policy anticipates, the policy usually ships a switch. A web application that must open outbound connections trips httpd_can_network_connect; a server that needs to send mail trips httpd_can_sendmail. getsebool -a | grep httpd lists them, and setsebool -P makes the change persistent. Booleans are reviewable and reversible in a way a custom module is not, so they come before audit2allow.
getsebool -a | grep -E "^httpd_can_network_connect "httpd_can_network_connect --> offsetsebool -P httpd_can_network_connect onsystemctl restart nginx && curl -sS -o /dev/null -w "%{http_code}\n" http://localhost/api/health200ausearch -m AVC -ts recent | wc -l0Debugging without disarming the host
setenforce 0 puts every domain on the host into permissive mode: all denials are logged and none is enforced, until someone remembers. When a service produces several denials and you want to see all of them at once instead of fixing one per restart, put that domain into permissive mode with semanage permissive -a httpd_t, collect the AVCs, fix labels and booleans, and remove the exception with semanage permissive -d httpd_t. The rest of the system stays enforcing. Only after labels and booleans are exhausted is audit2allow -M myapp the right tool, and the module it generates is a policy change to review like code, not a workaround to install and forget.
Which fix, by what the denial says
| Denial pattern | Meaning | Fix |
|---|---|---|
tcontext type is wrong for the path (user_home_t, default_t, tmp_t on real content) | the file was copied, moved or created outside the policy’s expectation | semanage fcontext -a + restorecon |
| tcontext is right, a network or mail permission is denied | a behaviour the policy anticipates but disables | setsebool -P <boolean> on |
| the service writes to a path the policy never expected | a genuinely new access pattern | a reviewed audit2allow -M module, or a path the policy already allows |
denials on /proc, /sys or another process | the program is doing something a confined domain should not | fix the program, or a different domain; not a module |
On OpenShift: SCCs, MCS labels and volumes
Containers get SELinux labels too. On OpenShift a Security Context Constraint assigns each namespace a pair of MCS categories (s0:c123,c456), so two containers on the same node run with different labels and cannot read each other's files even when both are root inside. The consequence for storage is that a volume must carry the Pod's label before the process can use it. Kubernetes handles that relabel at mount time: since v1.36, with seLinuxChangePolicy unset (the default, MountOption), the kubelet mounts eligible volumes with -o context=<label>, which is instant regardless of the volume's size. Setting Recursive restores the older behaviour of walking the filesystem and relabelling every file, which is slow on large volumes but is required when two Pods with different labels must share one volume on the same node; with the default, the second such Pod stays in ContainerCreating with a conflicting SELinux labels error.
spec:template:spec:securityContext:seLinuxOptions:level: "s0:c123,c456" # normally assigned by the SCC; set explicitly only to share a volume# seLinuxChangePolicy: Recursive # only when Pods with different labels share this volumecontainers:- name: appvolumeMounts:- name: datamountPath: /var/lib/appvolumes:- name: datapersistentVolumeClaim:claimName: app-data
The :Z and :z suffixes many guides mention are Docker and Podman volume options, not Kubernetes ones; a volumeMount has no such field. On a Pod, the label comes from seLinuxOptions (usually set by the SCC) and the relabel behaviour from seLinuxChangePolicy. When a mounted volume throws permission errors, oc get events for the Pod and ls -Z on the node's mount point say whether the label was applied, and the fix is the policy field above rather than a broader SCC.