Processes and /proc
ps, top, states, parents and /proc.
Every program that runs on a server is a process, and most questions about a busy or misbehaving machine start with a process: what is using the CPU, who started it, which service it belongs to, what it has open. This lesson shows how to answer those questions with ps, pstree, top and the /proc directory, how to read the state letters you will see in their output, and how to tell a routine "(deleted)" after a package upgrade from one that deserves a closer look.
Programs, processes, ps and top
A program is a file on disk, such as /usr/bin/python3; a process is one running copy of it: the kernel gives it its own memory, a process ID (PID) that no other running process has, the user it runs as, and a record of which process started it. One program can run as many processes at once, and a process disappears when it exits while the file stays.
In the example for this lesson, the server web01 feels slow. Another administrator, ess-processes-alice, is logged in over SSH and has started a script, report.py, that keeps one CPU busy. To reproduce it, open a second SSH login to your practice machine, save the script below as report.py, make it executable with chmod +x report.py and start it with ./report.py; your listings will show your own account name where these show alice's. It stops by itself after four minutes.
#!/usr/bin/python3# Stand-in for a colleague's forgotten job: keeps one CPU busy, appends a line to report.log now# and then, and stops by itself after four minutes.import timelog = open("report.log", "a")end = time.time() + 240n = 0while time.time() < end:n += 1if n % 20_000_000 == 0:log.write(f"{n}\n")log.flush()
The first command, typed in your own session, lists every process with the heaviest CPU users first:
ps prints a snapshot of the process table. The letters aux are old BSD-style options, written without a dash: a and x together select every process, including those with no terminal, and u picks the user-oriented columns. --sort=-%cpu sorts by CPU use, highest first, and head -n 5 keeps the header and the first four lines.
Read the columns from the left. USER is the account the process runs as; names longer than the column are cut short and end in +, so ess-pro+ is alice. %CPU is the CPU time the process has used divided by how long it has existed, so a value near 100 means it has kept one CPU busy for its whole life.
VSZ is the size of the memory the process has mapped, much of which it never uses, and RSS is the part actually held in RAM now, both in KiB; RSS is the one to watch. TTY names the terminal a process is attached to, pts/0 for an SSH login and ? for a background service. STAT is the process state, explained further down. TIME is the CPU time used so far, and COMMAND shows how it was started. The top line is alice's script. Below it are systemd (PID 1, busy after a recent boot), alice's per-user service manager systemd --user, and the agent of the lab's virtual machine tool.
ps is a snapshot; top redraws every few seconds with the busiest processes first, which is what you want while something is happening. Run top on its own in a terminal. The lab used batch mode (-b -n 1, one screen printed as text) so the output can be shown here:
The first five lines, left out here, summarise the whole machine and are the subject of the lesson on system resources. In the process list %CPU means something different from ps: it is the share of one CPU used since the previous refresh, so 100.0 is one CPU fully busy right now. RES is the same as RSS, S is the state and TIME+ the CPU time used. While top runs, P sorts by CPU, M by memory and q quits.
Who started it: parents and children
Every process except the first is started by another process, its parent. A parent creates a child by copying itself (the fork system call), and the copy then loads the new program (exec). The child records its parent's PID as its PPID. To see alice's processes with the columns you choose, give ps a list after -o and a user after -u:
ELAPSED is how long each process has been running. The PPID column links the lines together: report.py was started by the bash above it, which was started by sshd-session. To see the whole line of ancestors at once, pgrep finds the PID from the name and pstree walks up from it; -a adds arguments, -p PIDs and -s the parents:
Read the tree from the top. PID 1 is systemd, the first process the kernel starts. sshd is the SSH server's listener, the process that accepts connections on port 22 (in ps it shows as sshd: /usr/sbin/sshd -D [listener]). Since OpenSSH 9.8 the listener hands each new connection to a separate program, sshd-session, which appears twice: the first copy runs as root and does the privileged work such as checking the key, and the second runs as alice and owns her terminal. Then comes the shell that ran the command alice gave ssh, and finally the script. So this process came from an interactive login, not from a service or a scheduled job, and the person to ask about it is alice.
What /proc knows about a process
ps and top get their data from /proc, a filesystem the kernel generates on the fly: nothing in it is stored on disk, and every running process has a directory named after its PID. You can read those files yourself. status holds the process's identity in plain text:
Name is the short command name the kernel keeps (the script's file name here), State the current state and PPid the parent. Uid shows four numbers: the real, effective, saved and filesystem user IDs, which differ only for programs that change identity, such as sudo. 4161 is alice's user ID. The full command line is in cmdline, with the arguments separated by NUL bytes rather than spaces, which is why tr swaps them for spaces:
Anyone can read status and cmdline. The links that show which program file is running, the working directory and the open files are protected, because they reveal what another user is working on: the kernel lets only the process's owner and root follow them.
exe points to the file the kernel is running. For a script that is the interpreter, /usr/bin/python3.14, and the script's name is in cmdline. cwd is the directory the process works in. fd lists the open file descriptors, the numbered handles a process uses for everything it has open: 0, 1 and 2 are its input, output and error, here alice's terminal, and 3 is a log file open for writing (l-wx). A network connection would appear as socket:[...]. These few files answer most of the questions you would ask about an unfamiliar process.
Which service a process belongs to
systemd, the first process, starts and tracks everything else, and it keeps its records in units: a .service unit is a service such as cron, a .scope unit is a group of processes it did not start itself, such as one login session. It puts each unit's processes in their own control group, a kernel grouping of processes, so each process belongs to exactly one unit. The lesson "Services with systemd" covers units properly; here you only need to find them. ps can print the unit, and systemctl status accepts a PID as well as a unit name:
PID 1 sits in init.scope. The cron daemon belongs to cron.service, so it is managed with systemctl and logs to that unit's journal. Alice's script is in session-27.scope, the unit systemd creates for one login session; the number grows with every session opened since boot. When a process belongs to a service, stop or restart the service rather than killing the process, or systemd may simply start it again.
Process states
The STAT column (and State in /proc) says what a process is doing. The first letter is the state:
R is running or ready to run, waiting only for a CPU. S is sleeping until something happens, such as a key press, a network packet or a timer; most processes spend most of their time here. D is uninterruptible sleep, usually waiting for a disk or network storage to answer; it cannot be interrupted until the I/O finishes.
T is stopped, by Ctrl-Z or a stop signal (the next lesson). Z is a zombie: a process that has exited but whose parent has not yet collected its exit status. I marks idle kernel threads. Letters after the first add detail: s marks a session leader such as a login shell, and + a process in the foreground of its terminal, which is why alice's script showed R+. Counting the first letters over the whole machine gives a quick picture:
Most processes sleep, the I lines are kernel threads, and the R lines are the processes on a CPU at that moment (alice's script, ps itself). The one zombie was left by the lab virtual machine's own SSH connection, not by Ubuntu. A zombie uses no CPU and no memory, only an entry in the process table, and it cannot be killed because it has already exited. It disappears when its parent collects it, or when the parent exits and PID 1 adopts it. One zombie is harmless. A count that keeps growing means a parent program that never collects its children, and the fix is in that parent. A pile of D processes is a different signal: something they are all waiting for, usually storage, is slow or stuck.
"(deleted)" after an upgrade
Sooner or later ls -l /proc/PID/exe will show a path followed by (deleted). It means the file the process was started from no longer exists under that name, while the process keeps running the copy it loaded; the kernel keeps the old file's data until the last process using it exits. On a server the usual cause is a package upgrade. dpkg and rpm write the new version of a file under a temporary name and rename it over the old one, so the name now points to the new file and the running process holds the old, unnamed one.
To show this, the lab runs a small service from /usr/local/bin/ess-processes-demo (a copy of sleep) and then replaces that file the same way dpkg does. systemctl show -P MainPID prints the service's main PID:
The process is unchanged and still running the old version, which is exactly the problem after a security update: the fix is on disk, but the running service does not have it until it restarts. On Ubuntu Server, needrestart is installed for this and runs after every apt install or upgrade. Run by hand with -r l (list only), it restarts nothing and reports what should be restarted:
(The debconf lines left out at the top are a complaint that the lab had no terminal to show a menu in.) Only the demo service is listed, because the lab machine had just rebooted. A service still using a shared library that an upgrade replaced is listed as well; libraries show up as "(deleted)" in /proc/PID/maps rather than in exe.
When needrestart runs from apt, it does more than report. Ubuntu's default restart mode for the apt hook is automatic (the comments in /etc/needrestart/needrestart.conf say so): after apt installs or upgrades anything, from a script, from unattended-upgrades or at your prompt, it restarts the affected services itself and prints Restarting services.... A database using a replaced library is restarted without a question; the hardening lesson on patching shows how to change that. It defers a few services, such as D-Bus, whose restart can disrupt the whole machine. Restarting the demo gives the process the new file:
On RHEL the equivalent check is dnf needs-restarting. A "(deleted)" is worth a closer look when there was no recent upgrade (/var/log/apt/history.log on Ubuntu, dnf history on RHEL), when the path is not one that packages install to (dpkg -S PATH or rpm -qf PATH name the owning package), or when the file ran from a place such as /tmp, /var/tmp, /dev/shm or a home directory. Malware sometimes deletes its own file after starting; the advanced security course covers that case.
Try this
Copy a program and run the copy in the background: mkdir -p ~/processes and cp /usr/bin/gnusleep ~/processes/mysleep (the GNU sleep; Ubuntu's own sleep is a link into a multi-call binary that fails under another name), then ~/processes/mysleep 300 > /dev/null 2>&1 &; the & starts it in the background and $! holds its PID, both covered in the next lesson. Compare echo $$ (your shell's PID) with the PPid line of /proc/$!/status: they are the same, and the state is S (sleeping). ps -o pid,unit,cmd -p $! shows it in your login's session-N.scope. Now replace the file the way an upgrade would, with cp /usr/bin/gnusleep ~/processes/mysleep.new and mv ~/processes/mysleep.new ~/processes/mysleep, and look at ls -l /proc/$(pgrep mysleep)/exe: it ends in "(deleted)". Finish with pkill -x mysleep.
Takeaway
Before you kill a process you do not recognise, find out whose it is: pstree -aps PID for how it started, /proc/PID for what it runs and touches, and ps -o unit for the service it belongs to. A "(deleted)" executable after an upgrade usually means "restart me", so check needrestart and the package history before assuming the worst.