Signals and job control
kill, TERM vs KILL, jobs, fg, bg and nohup.
Stopping a program, pausing it, telling a daemon to reload, or pressing Ctrl-C all come down to one mechanism: the kernel delivering a signal to a process. This lesson shows which signals you will send, how a program can handle them and why that makes the order TERM first, KILL last, how to aim kill and pkill without hitting the wrong process, and how the shell runs jobs in the background. It ends with the practical question behind most of this: how to start long work over SSH so that it survives a dropped connection.
What a signal is
A signal is a small numbered notification that the kernel delivers to a process. Another process can send one with kill, the terminal sends one when you press Ctrl-C, and the kernel sends some itself. Every signal has a default action, usually ending the process, and a program can install a handler, a function of its own that runs instead, or choose to ignore the signal. Two signals are exceptions: KILL and STOP are carried out by the kernel and never reach the program's code. kill -l lists them all, and translates between numbers and names:
A handful of them cover daily work. TERM (15) is the polite request to exit, and it is what kill sends when you name no signal. KILL (9) makes the kernel end the process at once, with no chance to clean up. INT (2) is what Ctrl-C sends to the program in the foreground. HUP (1), "hangup", tells a process its terminal has gone away; many daemons treat it as "read your configuration again" instead. STOP (19) pauses a process and CONT (18) resumes it, and TSTP (20) is the pause that Ctrl-Z sends, which a program may handle. Use the names: the numbers above are the same on x86_64 and ARM servers, but some architectures number several signals differently (signal(7) has the table).
TERM first, KILL last
The difference between TERM and KILL is whether the program gets to clean up. To see it, here is a small script that stands in for a program holding a lock file, the kind of marker a program uses to say "I am running". Its trap line tells bash to run the cleanup function when TERM arrives. Create the directory with mkdir ~/signals, save the script there as worker.sh with any editor, and make it executable with chmod +x ~/signals/worker.sh:
#!/bin/bash# Stands in for a program that must clean up before it exits.lock=~/signals/worker.lockcleanup() {echo "TERM received: removing the lock and exiting"rm -f "$lock"exit 0}trap cleanup TERMtouch "$lock"echo "worker $$ started, lock taken"while true; do sleep 1; done
Start it in the background with &, so the shell carries on while it runs, and send it TERM. $! holds the PID of the last command started in the background, sleep 1 gives the script a moment to set its trap, and wait waits for the process to end and returns its exit status:
The handler ran, removed the lock and exited with the status the script chose, 0. Now the same with KILL:
No handler ran, and the lock file is still there. bash reports the job as Killed and its status as 137. When a process ends because of a signal, the shell reports 128 plus the signal number, so 137 means KILL, 143 means TERM and 130 means INT. A program that does not handle TERM gets the default action and ends with 143:
Those numbers are worth recognising in logs. status=9/KILL or an exit code of 137 from a service or a container means something forced it: a person, systemd's stop timeout, or the kernel's out-of-memory killer. An exit after TERM means whatever the program decided, which is why the first script reported 0.
Aiming kill and pkill
kill takes PIDs. pgrep finds PIDs by name and pkill sends a signal to every process it matches. Both treat the name as a pattern that may match part of a process name, which catches people out. In the lab a second account, ess-signals-bob, is logged in over SSH and has started an ssh-agent (the program that holds SSH keys, covered in "Connecting with SSH"). Suppose you want to end that agent. -u USER limits the match to one user's processes and -a prints each command line:
The pattern ssh matched the agent and also bob's login session, sshd-session: ess-signals-bob@notty (notty because this login runs a command without a terminal). sudo pkill -u ess-signals-bob ssh would have ended the agent and logged bob out. Without -u, the same pattern also matches the SSH server's listener (sshd: /usr/sbin/sshd -D [listener]) and every other user's session, yours included. -x matches the exact name, so the second command finds only the agent; -f matches against the whole command line instead of the name. Run pgrep -a with the same options first, read the list, and only then run pkill.
The kernel also checks who is asking. An ordinary user may signal only their own processes; root may signal any process.
(At an interactive prompt the message starts with -bash: kill:; the line 1 appears because the lab ran the command from a script.)
Sometimes you want a process to stop doing anything without ending it, for example a runaway job you want to look at before deciding, or a process you do not trust and want to examine first. STOP pauses it and CONT lets it continue:
T is the stopped state, and S shows it sleeping normally again after CONT. (The sleep 0.5 stands in for the moment you take to type the next command; without it, STOP can arrive before the new process has even started sleep.) A stopped process keeps its memory, open files and network connections, so programs on the other end of those connections may time out. For a whole service, sudo systemctl freeze UNIT pauses every process in it at once and systemctl status then reports active (running) (frozen) until sudo systemctl thaw UNIT.
D (uninterruptible sleep) is waiting inside the kernel, usually for a disk or a network filesystem, and signals are not delivered until that wait ends. kill -9 then appears to do nothing. Some of these waits do accept KILL, but most do not. The fix is to recover the storage or the mount it is waiting for; in the worst case, the machine has to be rebooted.Jobs: foreground and background
A job is a command line you started from your shell. Normally it runs in the foreground: the shell waits, and your keystrokes, including Ctrl-C, go to it. With & at the end it runs in the background and you get the prompt back at once. The shell numbers its jobs, and you can refer to them as %1, %2 and so on:
jobs -l lists the jobs with their PIDs. + marks the current job, the one that commands without a job number act on, and - the previous one. kill %1 sent TERM to job 1, and wait %1 returned its status, 143. In an interactive shell bash also prints the job number and PID when a background job starts ([1] 51792) and reports a finished job just before the next prompt ([1]+ Terminated followed by the command); the lab ran these commands without a terminal, where bash prints neither.
Three keys and two commands move jobs between the states. The messages quoted here come from an interactive bash that the lab ran in a pseudo-terminal, typing one key at a time. Ctrl-Z sends TSTP to the foreground job, which stops, and bash reports [1]+ Stopped sleep 300. bg resumes the stopped job in the background ([1]+ sleep 300 &), and fg brings a job back to the foreground, where Ctrl-C interrupts it and echo $? prints 130. If you type exit while a job is stopped, bash answers There are stopped jobs. and stays; a second exit leaves anyway.
Work that must outlive your SSH session
When an SSH connection drops or you close the terminal window, the kernel sends HUP to your login shell. An interactive bash then sends HUP to all of its jobs before it exits, and a job that does not handle HUP ends with it. The lab tested this with a second account: it logged in over SSH, started four background jobs, and then its SSH client was killed. The plain sleep 1001 & was gone afterwards; the three started in the ways below survived.
Killing the client closes the connection at once, so the server noticed immediately. When a network drops silently, as when a laptop loses Wi-Fi, the server only finds out when a keepalive or a write times out: sshd's own check (ClientAliveInterval) is off by default, and the kernel's TCP keepalive first probes after two hours. The hangup, and the end of the job, can therefore come long after the drop. Either way the job is at the mercy of the connection.
exit is different from a drop: bash sends HUP to its jobs on exit only if the huponexit option is set, which it is not by default, and systemd-logind does not kill a user's processes at logout unless KillUserProcesses=yes is configured (the default is no). A running background job survived exit in the lab. Do not rely on that: the connection drops when you least expect it.nohup starts a command with HUP ignored. Sending HUP to both jobs shows the difference:
The plain job ended with Hangup; the one under nohup is still running. At a terminal, nohup also says nohup: ignoring input and appending output to 'nohup.out': it sends the command's output to that file in the current directory, because the terminal it would have written to may disappear. For a job that is already running, disown removes it from the shell's job list, so bash will not send it HUP.
Both leave the job without a terminal you can come back to. tmux, a terminal multiplexer that Ubuntu Server installs (sudo dnf install tmux on RHEL), keeps a whole terminal session running on the server. Type tmux new -s work, start the job inside it, detach with Ctrl-b then d, and log out; later, tmux attach -t work puts you back in front of it. The tmux server is not a child of your login shell, so the hangup never reaches it. The same commands work without a terminal:
For real work on a server, such as a long data migration or a backup that must finish, run it as a service instead. A service is a program systemd, the first process on the machine, starts and supervises; systemd calls each thing it manages a unit, and the lesson "Services with systemd" covers them. For now you need three commands: systemctl status UNIT shows a unit's state, systemctl stop UNIT stops it, and journalctl -u UNIT shows what it logged (-I limits that to its latest run). systemd-run starts a command as a temporary service: it survives any logout, its output goes to the journal, and systemctl can show and stop it like any other service. --uid runs it as your user rather than as root:
systemctl stop does not send KILL straight away. It sends the unit's KillSignal to every process in the service, waits up to TimeoutStopSec, and only then sends FinalKillSignal. For most services that is TERM, 90 seconds, then KILL:
To watch the escalation, this service ignores TERM (trap "" TERM) and has its timeout cut to five seconds. time measures how long the stop takes, and journalctl -I shows only the service's most recent run:
systemd sent TERM at 08:35:23 (Stopping), waited five seconds, logged that the stop-sigterm state timed out, and sent KILL; the main process ended with status=9/KILL and the unit with the result timeout. When a service always takes exactly 90 seconds to stop, this is the journal pattern to look for: the program ignores or mishandles TERM.
Try this
Create ~/signals/worker.sh from this lesson and run it as a service: sudo systemd-run --uid=$USER --unit=worker -p TimeoutStopSec=5 $HOME/signals/worker.sh. Stop it with sudo systemctl stop worker, then read journalctl -u worker -I: the stop is immediate, the handler's message is in the journal and the lock file is gone. You will also see Terminated sleep 1, because systemd sends TERM to every process in the service, including the sleep the script was waiting for. Then copy the script to stubborn.sh, change its trap line to trap "" TERM, run and stop that the same way with time in front of the stop, and predict the journal before reading it: the stop takes about five seconds, the journal shows State 'stop-sigterm' timed out. Killing., and worker.lock is left behind.
Takeaway
Send TERM and give the program time; keep KILL for processes that no longer respond, and preview every pkill with pgrep -a. Anything that must survive your SSH session belongs in tmux or, better, in a service started with systemd-run.