/proc and /sys: the kernel's own view
Per-process status, stack, io and memory maps.
Every tool in the previous lesson is a formatter over files the kernel generates on demand under /proc and /sys. This lesson reads those files directly: a process's state, where it sleeps in the kernel, how much of its I/O reached the disk, how its memory is shared and how long it waited for a CPU, then the device tree behind a disk and a network interface. Knowing the source lets you tell what a tool rounds away, and it still works in a rescue shell or a minimal container image where the tool is not installed. The basics from Linux essentials (/proc/PID/cmdline, exe and fd, and ps itself) are assumed.
Files the tools are made of
/proc and /sys are pseudo-filesystems: nothing is stored on a disk, and each read runs kernel code that formats the current state as text. The system-wide files under /proc are where uptime and vmstat get their numbers.
/proc/loadavg holds the three load averages, then runnable tasks and all tasks as a fraction (3/161), then the last PID the kernel handed out. In /proc/stat, procs_running and procs_blocked are exactly vmstat's r and b columns. ctxt (context switches) and processes (tasks created) are counters since boot, so a tool reads them twice and divides the difference by the interval to get a rate.
A process with known behaviour
Numbers only mean something if you know what the process did, so the examples use a small Python program that writes 32 MiB, reads it back twice, holds 64 MiB of memory and then sleeps. Save it as /var/tmp/k-proc-demo and make it executable.
#!/usr/bin/python3# k-proc-demo: known I/O and memory use, then it waits so you can inspect it in /proc.import os, timepath = '/var/tmp/k-proc-demo.dat'with open(path, 'wb') as f:f.write(os.urandom(32 << 20)) # write 32 MiBf.flush()os.fsync(f.fileno()) # force it to the diskos.posix_fadvise(f.fileno(), 0, 0, os.POSIX_FADV_DONTNEED) # drop it from the page cachefor _ in range(2): # read it twice: first from disk, then from the cachewith open(path, 'rb') as f:f.read()ballast = b'x' * (64 << 20) # 64 MiB of private memory, every page touchedtime.sleep(3600)
posix_fadvise(... POSIX_FADV_DONTNEED) asks the kernel to drop the file's pages from the page cache, the memory the kernel uses to keep file data, so the first read has to go to the disk. Run the program as a transient service so it keeps running in the background under your own user.
Every process gets a directory like this, over fifty entries. pgrep -x matches the process name exactly. The name is k-proc-demo and not python3 because when the kernel runs a script through its #! line, it names the process after the script file.
State, and where a process sleeps
status is the readable summary that ps and top draw on. State: S (sleeping) means it waits for an event and can be woken by a signal. Tgid is the process ID; each thread has its own ID and shares this one, which the next lesson explains. PPid: 1 because systemd started it. VmRSS is resident memory, split into RssAnon (private memory such as the 64 MiB ballast) and RssFile (pages of files such as the Python binary and its libraries). Cpus_allowed_list is the CPUs it may run on.
The two context-switch counters say how the process gives up the CPU. A voluntary switch happens when it blocks to wait for I/O, a lock or a timer. A nonvoluntary switch happens when the scheduler takes the CPU away because its time is up or a more urgent task woke. A nonvoluntary count that climbs fast is a CPU-bound task competing for CPUs; a busy process with a mostly voluntary count is waiting on something else. This one was preempted about a dozen times while it generated and copied data in its first quarter of a second, and has hardly run since.
wchan (wait channel) names the kernel function a sleeping task is waiting in: here hrtimer_nanosleep, the timer behind Python's time.sleep(). For the full path through the kernel, read /proc/PID/stack.
The owner gets Permission denied: the kernel restricts stack dumps to root (CAP_SYS_ADMIN), because unwinding a running task's kernel stack can leak kernel memory. Read it bottom to top: the program entered the kernel with a system call (el0_svc is the aarch64 system-call entry from user mode; an x86_64 server shows entry_SYSCALL_64_after_hwframe and do_syscall_64 instead), called clock_nanosleep, and is waiting in hrtimer_nanosleep. For a task stuck in uninterruptible sleep, this file is what tells you which subsystem it is stuck in; the next lesson uses it for that.
I/O, memory and CPU wait, per process
rchar and wchar count bytes passed through read and write system calls, whether or not a disk was involved. read_bytes and write_bytes count bytes that this process caused to be fetched from or sent to the storage layer. The program read its 32 MiB file twice, so rchar is 64 MiB plus the Python modules it loaded, but read_bytes is exactly 32 MiB: the first read came from the disk and the second from the page cache. write_bytes is counted when pages are dirtied, before writeback: a write lands in the page cache first and is marked dirty, and the kernel writes it to the disk later, which is called writeback. When rchar is far larger than read_bytes, the page cache is doing its job; a process whose read_bytes climbs with every run is one that the cache cannot hold. syscr and syscw count the system calls: large reads in few calls here.
Not every file in /proc/PID is world-readable. status, stat and schedstat are. io, environ, fd and smaps_rollup pass the kernel's ptrace read-access check: you can read them for processes running under your own user, and root can read all of them. That check is weaker than the one for attaching a debugger, which Ubuntu restricts further with Yama's ptrace_scope setting (the strace and debugging lessons show that failure). stack is root-only, as above.
smaps_rollup sums the per-mapping memory counters in /proc/PID/smaps. Rss counts every page in RAM that the process maps. Pss (proportional set size) divides each shared page by the number of processes mapping it, so Pss_File is much smaller than Shared_Clean: shared libraries such as the C library are mapped by almost every process on the machine. Adding up Pss across processes approximates the memory they really use; adding up Rss counts every shared page many times. Swap is zero because this machine has no swap. To produce these sums the kernel walks every mapping of the process: fine once by hand, but do not poll smaps_rollup from monitoring on large processes such as databases or JVMs, where VmRSS and RssAnon in status are cheap to read.
AnonHugePages says that 67584 kB, the 64 MiB ballast included, sits in transparent huge pages (THP): 2 MiB pages the kernel can use instead of 4 KiB ones, which saves page-table work for large memory areas. Whether the kernel may do that is a system-wide mode in /sys:
The brackets mark the active mode. Ubuntu uses madvise: huge pages only for memory a program has marked with the madvise(MADV_HUGEPAGE) system call. The Python program never asks, but its C library does on its behalf: since glibc 2.43, malloc marks large allocations that way by default on aarch64 (the glibc release notes compare it to setting the glibc.malloc.hugetlb=1 tunable). On an x86_64 Ubuntu server, where glibc does not do this by default, the same ballast stays in 4 KiB pages. RHEL 10's kernel uses always, so large anonymous memory can get huge pages there without asking. The tuning lesson covers when to change the mode.
schedstat holds three numbers: nanoseconds spent running on a CPU, nanoseconds spent runnable but waiting for a CPU, and the number of times it was scheduled in. This process has mostly slept, so its wait is small. Under the CPU saturation of the previous lesson, the second number grows as fast as the first, and that is what pidstat reports as %wait.
The service has three file descriptors: standard input is /dev/null, and standard output and error are one socket, the connection systemd gives a service to the journal. limits shows a soft limit of 1024 open files and a hard limit of 524288, the systemd defaults for services; a process can raise its own soft limit up to the hard one. cgroup names the unit that owns the process, which is how systemctl status PID finds it.
/sys: the device model
/sys is where the kernel exposes devices, drivers and their settings. The real tree is /sys/devices, laid out by how the hardware is connected. /sys/block and /sys/class/net are indexes into it by type, made of symbolic links.
The disk vda and the interface eth0 both sit on virtio devices on the virtual PCI bus. On physical hardware the same links lead through a SATA, NVMe or network controller, which is how you find the driver and the slot behind a name.
The queue directory holds the block layer's settings for the disk. scheduler lists the I/O schedulers available, the kernel code that orders and merges requests before they reach the device, and brackets the active one: mq-deadline here, with none the alternative. RHEL 10 builds kyber and bfq into its kernel and lists them too; Ubuntu has them as modules that are not loaded, so they appear only after modprobe. The block-layer lesson compares them. rotational is what the driver reports, not a measurement: this virtual disk reports 1, a spinning disk, although it is a file on the host's SSD. stat holds the disk's counters since boot; the first four fields are reads completed, reads merged, sectors read and the total milliseconds read requests waited, and the same four for writes follow; the kernel's block-layer stat documentation describes the remaining fields. iostat reads the same numbers for every disk from /proc/diskstats and turns their differences into rates.
Each interface has one file per attribute: its state, its MTU and, under statistics, the same byte, packet, error and drop counters that ip -s link shows. The CPU example shows why you read the kernel's view through a tool that knows the platform: on this aarch64 machine /proc/cpuinfo has no model name line at all, because the Arm kernel reports implementer and part numbers instead. An x86_64 server prints a line such as model name : Intel(R) Xeon(R) ... for every CPU. lscpu reads /proc/cpuinfo and /sys/devices/system/cpu and gives one layout on both architectures.
/proc/sys and sysctl
/proc/sys holds the kernel's tunable parameters, and sysctl is a front end to it: a sysctl name is the path under /proc/sys with dots instead of slashes, so sysctl vm.swappiness and the command below read the same file.
vm.swappiness takes a value from 0 to 200 and sets the relative cost the kernel assumes for swapping anonymous memory against dropping file pages; 60 is the default. Values written to /proc/sys last until reboot. At boot, systemd-sysctl.service applies the .conf files from /usr/lib/sysctl.d (the distribution's defaults) and /etc/sysctl.d (yours). Ubuntu 26.04 has no /etc/sysctl.conf at all, so on Ubuntu a guide that tells you to edit it is out of date. RHEL 10 still ships the file and applies it through the link /etc/sysctl.d/99-sysctl.conf, but a named file in /etc/sysctl.d is the better place there too. The essentials lesson on the filesystem hierarchy showed the /usr/lib against /etc rule these directories follow; the hardening course's sysctl lesson (optional) goes into the full override order, and changing values for performance is the last lesson of this course, after you can measure the effect.
/proc/sys or /sys changes the running kernel with no confirmation and no undo. Read the current value first, change one setting at a time, and measure before you make it permanent.Try this
Predict before you look. Stop the demo with sudo systemctl stop k-proc-demo, delete the posix_fadvise line from the program, start it again with the same systemd-run command and read its io file. Without that line the freshly written file stays in the page cache, so both reads are cache hits: rchar is still about 64 MiB, write_bytes is still 32 MiB, and read_bytes is 0. Then compare VmRSS in status with Pss in smaps_rollup, and finish with sudo systemctl stop k-proc-demo and rm /var/tmp/k-proc-demo /var/tmp/k-proc-demo.dat.
Takeaway
When a tool's number surprises you, read the file it came from: status, io, smaps_rollup and schedstat under /proc/PID separate what the process asked for from what the kernel actually did.