/proc and /sys: the kernel's own view

Per-process status, stack, io and memory maps.

Advanced14 min · lesson 2 of 21

Every tool in the previous lesson is a formatter over files the kernel generates on demand under /proc and /sys. This lesson reads those files directly: a process's state, where it sleeps in the kernel, how much of its I/O reached the disk, how its memory is shared and how long it waited for a CPU, then the device tree behind a disk and a network interface. Knowing the source lets you tell what a tool rounds away, and it still works in a rescue shell or a minimal container image where the tool is not installed. The basics from Linux essentials (/proc/PID/cmdline, exe and fd, and ps itself) are assumed.

Files the tools are made of

/proc and /sys are pseudo-filesystems: nothing is stored on a disk, and each read runs kernel code that formats the current state as text. The system-wide files under /proc are where uptime and vmstat get their numbers.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/loadavg
2.89 2.68 1.34 3/161 245823
$ grep -E '^(ctxt|processes|procs_running|procs_blocked)' /proc/stat
ctxt 6453755 processes 245861 procs_running 1 procs_blocked 0

/proc/loadavg holds the three load averages, then runnable tasks and all tasks as a fraction (3/161), then the last PID the kernel handed out. In /proc/stat, procs_running and procs_blocked are exactly vmstat's r and b columns. ctxt (context switches) and processes (tasks created) are counters since boot, so a tool reads them twice and divides the difference by the interval to get a rate.

Which file each tool reads
/proc/PID
status, stat
ps and top
io
pidstat -d
schedstat
pidstat %wait
/proc
stat
vmstat r, b and cpu columns
loadavg
uptime
pressure/*
PSI, read directly
/sys
block/DEV/stat
per-disk counters, as in /proc/diskstats
class/net/IF/statistics
the counters ip -s link shows
devices/system/cpu
lscpu
iostat reads /proc/diskstats, which holds the same fields for every disk.

A process with known behaviour

Numbers only mean something if you know what the process did, so the examples use a small Python program that writes 32 MiB, reads it back twice, holds 64 MiB of memory and then sleeps. Save it as /var/tmp/k-proc-demo and make it executable.

/var/tmp/k-proc-demo
#!/usr/bin/python3
# k-proc-demo: known I/O and memory use, then it waits so you can inspect it in /proc.
import os, time
path = '/var/tmp/k-proc-demo.dat'
with open(path, 'wb') as f:
f.write(os.urandom(32 << 20)) # write 32 MiB
f.flush()
os.fsync(f.fileno()) # force it to the disk
os.posix_fadvise(f.fileno(), 0, 0, os.POSIX_FADV_DONTNEED) # drop it from the page cache
for _ in range(2): # read it twice: first from disk, then from the cache
with open(path, 'rb') as f:
f.read()
ballast = b'x' * (64 << 20) # 64 MiB of private memory, every page touched
time.sleep(3600)

posix_fadvise(... POSIX_FADV_DONTNEED) asks the kernel to drop the file's pages from the page cache, the memory the kernel uses to keep file data, so the first read has to go to the disk. Run the program as a transient service so it keeps running in the background under your own user.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --unit=k-proc-demo --uid=$USER /var/tmp/k-proc-demo
Running as unit: k-proc-demo.service; invocation ID: f25f2158796440cb8ca386aabd6c4afb
$ ls -C /proc/$(pgrep -x k-proc-demo)
attr fd mem patch_state stat autogroup fdinfo mountinfo personality statm auxv gid_map mounts projid_map status cgroup io mountstats root syscall clear_refs ksm_merging_pages net sched task cmdline ksm_stat ns schedstat timens_offsets comm latency numa_maps sessionid timers coredump_filter limits oom_adj setgroups timerslack_ns cwd loginuid oom_score smaps uid_map environ map_files oom_score_adj smaps_rollup wchan exe maps pagemap stack

Every process gets a directory like this, over fifty entries. pgrep -x matches the process name exactly. The name is k-proc-demo and not python3 because when the kernel runs a script through its #! line, it names the process after the script file.

State, and where a process sleeps

deploy@web01 · Ubuntu 26.04 LTS
$ grep -E '^(Name|State|Tgid|PPid|Uid|VmRSS|RssAnon|RssFile|Threads|Cpus_allowed_list|voluntary|nonvoluntary)' /proc/$(pgrep -x k-proc-demo)/status
Name: k-proc-demo State: S (sleeping) Tgid: 245898 PPid: 1 Uid: 1001 1001 1001 1001 VmRSS: 76772 kB RssAnon: 70836 kB RssFile: 5936 kB Threads: 1 Cpus_allowed_list: 0-1 voluntary_ctxt_switches: 6 nonvoluntary_ctxt_switches: 14

status is the readable summary that ps and top draw on. State: S (sleeping) means it waits for an event and can be woken by a signal. Tgid is the process ID; each thread has its own ID and shares this one, which the next lesson explains. PPid: 1 because systemd started it. VmRSS is resident memory, split into RssAnon (private memory such as the 64 MiB ballast) and RssFile (pages of files such as the Python binary and its libraries). Cpus_allowed_list is the CPUs it may run on.

The two context-switch counters say how the process gives up the CPU. A voluntary switch happens when it blocks to wait for I/O, a lock or a timer. A nonvoluntary switch happens when the scheduler takes the CPU away because its time is up or a more urgent task woke. A nonvoluntary count that climbs fast is a CPU-bound task competing for CPUs; a busy process with a mostly voluntary count is waiting on something else. This one was preempted about a dozen times while it generated and copied data in its first quarter of a second, and has hardly run since.

deploy@web01 · Ubuntu 26.04 LTS
$ ps -o pid,stat,wchan:32,comm -p $(pgrep -x k-proc-demo)
PID STAT WCHAN COMMAND 245898 Ss hrtimer_nanosleep k-proc-demo

wchan (wait channel) names the kernel function a sleeping task is waiting in: here hrtimer_nanosleep, the timer behind Python's time.sleep(). For the full path through the kernel, read /proc/PID/stack.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/$(pgrep -x k-proc-demo)/stack
cat: /proc/245898/stack: Permission denied
$ sudo cat /proc/$(pgrep -x k-proc-demo)/stack
[<0>] hrtimer_nanosleep+0x94/0x128 [<0>] common_nsleep_timens+0x58/0xc8 [<0>] __arm64_sys_clock_nanosleep+0xf0/0x1a8 [<0>] invoke_syscall.constprop.0+0x60/0xf0 [<0>] el0_svc_common.constprop.0+0x114/0x140 [<0>] do_el0_svc+0x28/0x58 [<0>] el0_svc+0x40/0x1e0 [<0>] el0t_64_sync_handler+0xc0/0x110 [<0>] el0t_64_sync+0x1b8/0x1c0

The owner gets Permission denied: the kernel restricts stack dumps to root (CAP_SYS_ADMIN), because unwinding a running task's kernel stack can leak kernel memory. Read it bottom to top: the program entered the kernel with a system call (el0_svc is the aarch64 system-call entry from user mode; an x86_64 server shows entry_SYSCALL_64_after_hwframe and do_syscall_64 instead), called clock_nanosleep, and is waiting in hrtimer_nanosleep. For a task stuck in uninterruptible sleep, this file is what tells you which subsystem it is stuck in; the next lesson uses it for that.

I/O, memory and CPU wait, per process

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/$(pgrep -x k-proc-demo)/io
rchar: 67213243 wchar: 33554525 syscr: 94 syscw: 5 read_bytes: 33554432 write_bytes: 33554432 cancelled_write_bytes: 0

rchar and wchar count bytes passed through read and write system calls, whether or not a disk was involved. read_bytes and write_bytes count bytes that this process caused to be fetched from or sent to the storage layer. The program read its 32 MiB file twice, so rchar is 64 MiB plus the Python modules it loaded, but read_bytes is exactly 32 MiB: the first read came from the disk and the second from the page cache. write_bytes is counted when pages are dirtied, before writeback: a write lands in the page cache first and is marked dirty, and the kernel writes it to the disk later, which is called writeback. When rchar is far larger than read_bytes, the page cache is doing its job; a process whose read_bytes climbs with every run is one that the cache cannot hold. syscr and syscw count the system calls: large reads in few calls here.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/1/io
cat: /proc/1/io: Permission denied

Not every file in /proc/PID is world-readable. status, stat and schedstat are. io, environ, fd and smaps_rollup pass the kernel's ptrace read-access check: you can read them for processes running under your own user, and root can read all of them. That check is weaker than the one for attaching a debugger, which Ubuntu restricts further with Yama's ptrace_scope setting (the strace and debugging lessons show that failure). stack is root-only, as above.

deploy@web01 · Ubuntu 26.04 LTS
$ grep -E '^(Rss|Pss|Pss_Anon|Pss_File|Shared_Clean|Private_Dirty|AnonHugePages|Swap):' /proc/$(pgrep -x k-proc-demo)/smaps_rollup
Rss: 76772 kB Pss: 72367 kB Pss_Anon: 70836 kB Pss_File: 1531 kB Shared_Clean: 5848 kB Private_Dirty: 70836 kB AnonHugePages: 67584 kB Swap: 0 kB

smaps_rollup sums the per-mapping memory counters in /proc/PID/smaps. Rss counts every page in RAM that the process maps. Pss (proportional set size) divides each shared page by the number of processes mapping it, so Pss_File is much smaller than Shared_Clean: shared libraries such as the C library are mapped by almost every process on the machine. Adding up Pss across processes approximates the memory they really use; adding up Rss counts every shared page many times. Swap is zero because this machine has no swap. To produce these sums the kernel walks every mapping of the process: fine once by hand, but do not poll smaps_rollup from monitoring on large processes such as databases or JVMs, where VmRSS and RssAnon in status are cheap to read.

AnonHugePages says that 67584 kB, the 64 MiB ballast included, sits in transparent huge pages (THP): 2 MiB pages the kernel can use instead of 4 KiB ones, which saves page-table work for large memory areas. Whether the kernel may do that is a system-wide mode in /sys:

deploy@web01 · Ubuntu 26.04 LTS
$ cat /sys/kernel/mm/transparent_hugepage/enabled
always [madvise] never

The brackets mark the active mode. Ubuntu uses madvise: huge pages only for memory a program has marked with the madvise(MADV_HUGEPAGE) system call. The Python program never asks, but its C library does on its behalf: since glibc 2.43, malloc marks large allocations that way by default on aarch64 (the glibc release notes compare it to setting the glibc.malloc.hugetlb=1 tunable). On an x86_64 Ubuntu server, where glibc does not do this by default, the same ballast stays in 4 KiB pages. RHEL 10's kernel uses always, so large anonymous memory can get huge pages there without asking. The tuning lesson covers when to change the mode.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/$(pgrep -x k-proc-demo)/schedstat
233032041 1527833 20

schedstat holds three numbers: nanoseconds spent running on a CPU, nanoseconds spent runnable but waiting for a CPU, and the number of times it was scheduled in. This process has mostly slept, so its wait is small. Under the CPU saturation of the previous lesson, the second number grows as fast as the first, and that is what pidstat reports as %wait.

deploy@web01 · Ubuntu 26.04 LTS
$ ls -l /proc/$(pgrep -x k-proc-demo)/fd
total 0 lr-x------ 1 deploy deploy 64 Sep 27 09:30 0 -> /dev/null lrwx------+ 1 deploy deploy 64 Sep 27 09:30 1 -> socket:[844182] lrwx------+ 1 deploy deploy 64 Sep 27 09:30 2 -> socket:[844182]
$ grep 'open files' /proc/$(pgrep -x k-proc-demo)/limits
Max open files 1024 524288 files
$ cat /proc/$(pgrep -x k-proc-demo)/cgroup
0::/system.slice/k-proc-demo.service

The service has three file descriptors: standard input is /dev/null, and standard output and error are one socket, the connection systemd gives a service to the journal. limits shows a soft limit of 1024 open files and a hard limit of 524288, the systemd defaults for services; a process can raise its own soft limit up to the hard one. cgroup names the unit that owns the process, which is how systemctl status PID finds it.

/sys: the device model

/sys is where the kernel exposes devices, drivers and their settings. The real tree is /sys/devices, laid out by how the hardware is connected. /sys/block and /sys/class/net are indexes into it by type, made of symbolic links.

deploy@web01 · Ubuntu 26.04 LTS
$ ls -l /sys/block/vda /sys/class/net/eth0
lrwxrwxrwx 1 root root 0 Sep 27 08:26 /sys/block/vda -> ../devices/pci0000:00/0000:00:06.0/virtio2/block/vda lrwxrwxrwx 1 root root 0 Sep 27 08:26 /sys/class/net/eth0 -> ../../devices/pci0000:00/0000:00:01.0/virtio0/net/eth0

The disk vda and the interface eth0 both sit on virtio devices on the virtual PCI bus. On physical hardware the same links lead through a SATA, NVMe or network controller, which is how you find the driver and the slot behind a name.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /sys/block/vda/queue/scheduler /sys/block/vda/queue/rotational
none [mq-deadline] 1
$ cat /sys/block/vda/stat
61734 15853 5069068 18402 61955 89824 12079694 92618 0 18064 115546 4960 0 57188016 959 13511 3567

The queue directory holds the block layer's settings for the disk. scheduler lists the I/O schedulers available, the kernel code that orders and merges requests before they reach the device, and brackets the active one: mq-deadline here, with none the alternative. RHEL 10 builds kyber and bfq into its kernel and lists them too; Ubuntu has them as modules that are not loaded, so they appear only after modprobe. The block-layer lesson compares them. rotational is what the driver reports, not a measurement: this virtual disk reports 1, a spinning disk, although it is a file on the host's SSD. stat holds the disk's counters since boot; the first four fields are reads completed, reads merged, sectors read and the total milliseconds read requests waited, and the same four for writes follow; the kernel's block-layer stat documentation describes the remaining fields. iostat reads the same numbers for every disk from /proc/diskstats and turns their differences into rates.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /sys/class/net/eth0/operstate /sys/class/net/eth0/mtu /sys/class/net/eth0/statistics/rx_bytes
up 1500 339166732
$ grep -c 'model name' /proc/cpuinfo lscpu | head -7
0 Architecture: aarch64 CPU op-mode(s): 64-bit Byte Order: Little Endian CPU(s): 2 On-line CPU(s) list: 0,1 Vendor ID: Apple Model name: -

Each interface has one file per attribute: its state, its MTU and, under statistics, the same byte, packet, error and drop counters that ip -s link shows. The CPU example shows why you read the kernel's view through a tool that knows the platform: on this aarch64 machine /proc/cpuinfo has no model name line at all, because the Arm kernel reports implementer and part numbers instead. An x86_64 server prints a line such as model name : Intel(R) Xeon(R) ... for every CPU. lscpu reads /proc/cpuinfo and /sys/devices/system/cpu and gives one layout on both architectures.

/proc/sys and sysctl

/proc/sys holds the kernel's tunable parameters, and sysctl is a front end to it: a sysctl name is the path under /proc/sys with dots instead of slashes, so sysctl vm.swappiness and the command below read the same file.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/sys/vm/swappiness
60
$ ls /usr/lib/sysctl.d/ /etc/sysctl.d/ /etc/sysctl.conf
ls: cannot access '/etc/sysctl.conf': No such file or directory /etc/sysctl.d/: 99-cloudimg-ipv6.conf 99-lima.conf README.sysctl /usr/lib/sysctl.d/: 10-apparmor.conf 10-coredump-debian.conf 50-default.conf 50-pid-max.conf 55-bufferbloat.conf 55-console-messages.conf 55-ipv6-privacy.conf 55-kernel-hardening.conf 55-magic-sysrq.conf 55-map-count.conf 55-network-security.conf 55-ptrace.conf 55-zeropage.conf

vm.swappiness takes a value from 0 to 200 and sets the relative cost the kernel assumes for swapping anonymous memory against dropping file pages; 60 is the default. Values written to /proc/sys last until reboot. At boot, systemd-sysctl.service applies the .conf files from /usr/lib/sysctl.d (the distribution's defaults) and /etc/sysctl.d (yours). Ubuntu 26.04 has no /etc/sysctl.conf at all, so on Ubuntu a guide that tells you to edit it is out of date. RHEL 10 still ships the file and applies it through the link /etc/sysctl.d/99-sysctl.conf, but a named file in /etc/sysctl.d is the better place there too. The essentials lesson on the filesystem hierarchy showed the /usr/lib against /etc rule these directories follow; the hardening course's sysctl lesson (optional) goes into the full override order, and changing values for performance is the last lesson of this course, after you can measure the effect.

A write is live immediately
Writing to /proc/sys or /sys changes the running kernel with no confirmation and no undo. Read the current value first, change one setting at a time, and measure before you make it permanent.

Try this

Predict before you look. Stop the demo with sudo systemctl stop k-proc-demo, delete the posix_fadvise line from the program, start it again with the same systemd-run command and read its io file. Without that line the freshly written file stays in the page cache, so both reads are cache hits: rchar is still about 64 MiB, write_bytes is still 32 MiB, and read_bytes is 0. Then compare VmRSS in status with Pss in smaps_rollup, and finish with sudo systemctl stop k-proc-demo and rm /var/tmp/k-proc-demo /var/tmp/k-proc-demo.dat.

Takeaway

When a tool's number surprises you, read the file it came from: status, io, smaps_rollup and schedstat under /proc/PID separate what the process asked for from what the kernel actually did.

Quick check
01A nightly backup job reads 200 GB. Its /proc/PID/io shows rchar at about 200 GB and read_bytes at about 3 GB. What does that tell you?
Incorrect — The counters in /proc/PID/io accumulate for the life of the process; nothing resets them.
Correct — rchar counts bytes passed through read calls; read_bytes counts bytes the process caused to be fetched from the storage layer.
Incorrect — rchar counts bytes returned by read-type system calls. Data the job sends out is written, so it would appear in wchar.
Incorrect — read_bytes counts what the kernel fetched from the storage layer for this process. Compression inside a device happens below that and does not change it.
02A service is slow. Its /proc/PID/status shows nonvoluntary_ctxt_switches rising by thousands per second while voluntary_ctxt_switches barely moves. What is the most likely explanation?
Incorrect — Blocking on I/O is the textbook voluntary switch: the task gives up the CPU itself because it cannot continue.
Incorrect — A major page fault makes the task block for I/O, which counts as voluntary, and this host would show swap activity.
Incorrect — The thread count alone does not decide how a task loses the CPU. Preemption is what drives the nonvoluntary counter.
Correct — The scheduler takes the CPU away when the task's time is up or another task must run, which is CPU contention.
03A developer can read /proc/PID/status and /proc/PID/io for their own process, but cat /proc/PID/stack returns "Permission denied". What is going on?
Correct — The kernel checks for CAP_SYS_ADMIN before unwinding a task's kernel stack, so even the owner needs sudo.
Incorrect — Lowering ptrace_scope would not help: the check that fails is a capability check for root, and changing a host-wide protection for this is the wrong trade.
Incorrect — The kernel reads the stack of a sleeping task without stopping it; being stopped does not change who may read it.
Incorrect — hidepid hides other users' process directories, and this developer can read the status file in the same directory.

Related