System resources: CPU, memory and logins
Load, memory, uptime and who is logged in.
When a server is slow or behaves oddly, the first minutes go on three questions: how busy is it, is it short of memory, and who else is on it. This lesson reads the answers from uptime, top, vmstat, free and the kernel's pressure files, shows how to see who is logged in on Ubuntu 26.04, where the familiar who and last no longer tell you, and checks the clock, because every log timestamp depends on it. These are first readings; the internals course measures CPU and memory in depth.
To give the tools something to show, the lab machine was running a load generator, stress-ng, with three workers that each keep one CPU busy, and it had been running for a minute when the readings below were taken. A second administrator, ess-cmd-monitor-alice, was logged in over SSH.
How busy: uptime and the load average
uptime prints the time, how long the machine has been running since its last boot, how many users are logged in, and three load averages. The load average is the average number of processes that were running on a CPU, waiting for one, or waiting in uninterruptible sleep (state D, usually for a disk), over the last 1, 5 and 15 minutes. It is not divided by the number of CPUs, so compare it with nproc, which prints how many CPUs this machine may use: 2. A load that stays above the CPU count means work is queueing. The three numbers also show direction: the one-minute figure is well above the other two, so the load rose recently, which is when stress-ng began. It is still climbing towards 3, the number of busy workers, because it is an average that takes a few minutes to catch up. Because D processes count too, a high load with idle CPUs points at storage rather than at the CPU.
The summary lines of top
The first five lines of top put the most important numbers on one screen:
The first line repeats uptime. Tasks counts processes by state; the one zombie belongs to the lab virtual machine's own SSH connection (the processes lesson explains zombies). %Cpu(s) splits all CPU time into kinds: us running programs, sy the kernel working for them, ni programs with lowered priority, id idle, wa idle while waiting for disk I/O, hi and si handling interrupts, and st time the hypervisor gave to other virtual machines, which matters on cloud servers. The ones to read first are us, sy, id and wa. Here id is 0.0, so the CPUs are fully used, and wa is 0.0, so nothing is waiting for the disk: the machine is short of CPU. The two memory lines are the same figures that free shows, covered below.
Is work waiting? vmstat and pressure
vmstat 1 5 prints one line per second, five times. The first line is an average since boot and is usually ignored; the others describe each second:
Start with three columns. r is the number of processes running or waiting for a CPU; with r above the CPU count every second, processes are queueing for CPU time. b is the number blocked waiting for I/O. si and so show swapping in and out, and anything above zero there for long means memory is short.
The rest can wait until you need them: the memory columns are in KiB, bi and bo count blocks read from and written to disk, in and cs interrupts and context switches per second, and the last columns are the same CPU split as in top. vmstat is a good first command on a machine you do not know, because it shows CPU, memory, swap and disk activity side by side as they change.
The kernel can also say how much time work spent waiting. Pressure stall information (PSI) lives in /proc/pressure, with a file each for cpu, memory and io:
some avg10=50.40 means that during the last 10 seconds, runnable processes were kept waiting for a CPU about half of the time; avg60 and avg300 cover one and five minutes, and total is the accumulated waiting time in microseconds. For the CPU, the full line (every process stalled at once) is not defined for the whole system and stays at zero. PSI answers "is anything being slowed down?" directly, where the load average only counts processes. It is available on Ubuntu 26.04; the RHEL 10 kernel is built with PSI switched off unless the machine boots with psi=1:
Once the lab stopped the load, the pressure figures fell within seconds while the load average took longer:
Twenty-five seconds later avg10 had dropped from 50 to 7, while the one-minute load average had only fallen from 2.23 to 1.70, because it forgets old values gradually. Use PSI or vmstat to see what is happening now and the load average to see the trend.
Memory: read available, not free
free -h shows memory in readable units. total is the RAM the kernel can use. free is memory nothing is using at all, and on a healthy server it is often small, because the kernel keeps recently read file data in RAM (buff/cache) instead of leaving memory empty; most of that cache is handed back as soon as a program needs the memory. available is the kernel's estimate of how much a new program could get without swapping, and it is the number to read: here 3.3 GiB of 3.8. In this version of free, used is simply total minus available. shared is mostly files in memory-backed filesystems such as /tmp. The Swap line is all zeros because cloud images such as this one come without swap. All of these come from /proc/meminfo, in KiB:
When available approaches zero, the kernel's out-of-memory killer ends a process to free memory, and a service can be killed the same way for exceeding its own memory limit (MemoryMax=) while the machine as a whole still has memory left. Either way the kernel logs the kill, so journalctl -k is where to look when a process vanished without a trace in its own log.
Who is logged in, and who was
On Ubuntu 26.04, who prints nothing, even with two users logged in:
The default who reads /run/utmp, a file of current logins that dates from long before systemd, and Ubuntu 26.04 no longer creates it: its record format stores times in 32 bits, which overflow in 2038. Its companion history file, /var/log/wtmp, still exists, and sshd still appends logins that have a terminal to it, but Ubuntu no longer installs the tools that read it (below). The current sessions are known to systemd-logind, the service that tracks logins. The GNU who, installed as gnuwho (the first lesson explained the gnu* commands), asks it and does list the sessions:
Two standard commands ask logind as well, with more detail:
w shows each user, the terminal (pts/0 is an SSH login with a terminal), where the login came from, when it started, how long the user has been idle, and what they are running. It cuts user names to eight characters, so alice appears as ess-cmd-. The lima line is the lab virtual machine tool's own connection. loginctl list-sessions lists the sessions with full user names; the manager sessions belong to each user's own service manager, not to extra logins.
Login history used to come from last (successful logins) and lastb (failed ones), which read /var/log/wtmp and /var/log/btmp. Ubuntu 26.04 no longer installs them:
Every SSH login is still recorded by the SSH server in the journal and in /var/log/auth.log, which the lesson "Logs: the journal and /var/log" covers in detail. To see who logged in, search for the Accepted lines:
Each line gives the account, the address it came from and the key that was used. If you want the last command back, the wtmpdb package in Ubuntu's universe repository provides it (as a link to wtmpdb), with a PAM module that records logins in a database that has no 2038 problem. RHEL 10 still ships who, last and lastb:
Is the clock right?
A wrong clock breaks more than timestamps. TLS certificates are valid only between two dates, one-time login codes depend on the time, and logs from several servers can only be lined up into one timeline if their clocks agree.
date prints local time and date -u UTC; this server's time zone is UTC, as on most servers, so they match. timedatectl adds the hardware clock (RTC), the time zone, and the two lines that matter: NTP service: active means a time synchronisation service is running, and System clock synchronized: yes means it has set the clock from its time servers. On Ubuntu 26.04 that service is chrony, and chronyc sources shows which servers it uses:
Each line is a time server, and ^ marks a server. The second column is chrony's verdict: * is the source the clock is synchronised to, + a source combined with it, - a usable source that is not currently selected, and ? one it cannot use. Reach is an octal record of the last eight polls: 377 means all eight were answered and 0 means none was. Here every server shows 377. When the clock drifts or timedatectl says no, this is the check to make: if every line shows ^? with Reach 0, no server answers, and the fault lies between the server and its time sources.
Ubuntu 26.04's chrony also uses NTS (Network Time Security, authenticated NTP), which needs outbound TCP port 4460 as well as UDP port 123; the Linux hardening course covers it in "Logs and clocks you can trust".
Try this
On your practice machine, start three busy loops in the background with yes > /dev/null & three times (yes prints a line endlessly, and the output is thrown away). Wait about fifteen seconds, then run uptime, vmstat 1 3 and cat /proc/pressure/cpu. On a two-CPU machine r shows three or more, id drops to 0 and some avg10 climbs well above zero. Most of the CPU time appears under sy rather than us, because yes spends it inside the kernel writing to /dev/null. Stop the loops with kill %1 %2 %3, confirm with pgrep -c -x yes (it prints 0), and watch avg10 fall over the next half minute while the load average takes several minutes.
Takeaway
Compare the load with the CPU count, read available for memory, and check wa and PSI before blaming the CPU. On Ubuntu 26.04, ask w or loginctl who is logged in and the journal who has been, and make sure timedatectl says the clock is synchronised before trusting any timeline.