System resources: CPU, memory and logins

Load, memory, uptime and who is logged in.

Beginner12 min · lesson 18 of 29

When a server is slow or behaves oddly, the first minutes go on three questions: how busy is it, is it short of memory, and who else is on it. This lesson reads the answers from uptime, top, vmstat, free and the kernel's pressure files, shows how to see who is logged in on Ubuntu 26.04, where the familiar who and last no longer tell you, and checks the clock, because every log timestamp depends on it. These are first readings; the internals course measures CPU and memory in depth.

To give the tools something to show, the lab machine was running a load generator, stress-ng, with three workers that each keep one CPU busy, and it had been running for a minute when the readings below were taken. A second administrator, ess-cmd-monitor-alice, was logged in over SSH.

How busy: uptime and the load average

deploy@web01 · Ubuntu 26.04 LTS
$ uptime
11:13:01 up 49 min, 2 users, load average: 2.23, 1.23, 0.88
$ nproc
2

uptime prints the time, how long the machine has been running since its last boot, how many users are logged in, and three load averages. The load average is the average number of processes that were running on a CPU, waiting for one, or waiting in uninterruptible sleep (state D, usually for a disk), over the last 1, 5 and 15 minutes. It is not divided by the number of CPUs, so compare it with nproc, which prints how many CPUs this machine may use: 2. A load that stays above the CPU count means work is queueing. The three numbers also show direction: the one-minute figure is well above the other two, so the load rose recently, which is when stress-ng began. It is still climbing towards 3, the number of busy workers, because it is an average that takes a few minutes to catch up. Because D processes count too, a high load with idle CPUs points at storage rather than at the CPU.

The summary lines of top

The first five lines of top put the most important numbers on one screen:

deploy@web01 · Ubuntu 26.04 LTS
$ top -b -n 1 | head -n 5
top - 11:13:01 up 49 min, 2 users, load average: 2.23, 1.23, 0.88 Tasks: 136 total, 4 running, 131 sleeping, 0 stopped, 1 zombie %Cpu(s):100.0 us, 0.0 sy, 0.0 ni, 0.0 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st MiB Mem : 3895.8 total, 766.5 free, 539.4 used, 2780.5 buff/cache MiB Swap: 0.0 total, 0.0 free, 0.0 used. 3356.4 avail Mem

The first line repeats uptime. Tasks counts processes by state; the one zombie belongs to the lab virtual machine's own SSH connection (the processes lesson explains zombies). %Cpu(s) splits all CPU time into kinds: us running programs, sy the kernel working for them, ni programs with lowered priority, id idle, wa idle while waiting for disk I/O, hi and si handling interrupts, and st time the hypervisor gave to other virtual machines, which matters on cloud servers. The ones to read first are us, sy, id and wa. Here id is 0.0, so the CPUs are fully used, and wa is 0.0, so nothing is waiting for the disk: the machine is short of CPU. The two memory lines are the same figures that free shows, covered below.

Is work waiting? vmstat and pressure

vmstat 1 5 prints one line per second, five times. The first line is an average since boot and is usually ignored; the others describe each second:

deploy@web01 · Ubuntu 26.04 LTS
$ vmstat 1 5
procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu------- r b swpd free buff cache si so bi bo in cs us sy id wa st gu 4 0 0 784712 151868 2695348 0 0 2810 1667 1410 9 18 9 73 1 0 0 3 0 0 784616 151868 2695352 0 0 0 0 2006 562 100 0 0 0 0 0 3 0 0 784616 151868 2695352 0 0 0 0 2007 537 100 0 0 0 0 0 3 0 0 784616 151868 2695352 0 0 0 80 2009 546 100 1 0 0 0 0 3 0 0 784616 151868 2695352 0 0 0 0 2008 539 100 0 0 0 0 0

Start with three columns. r is the number of processes running or waiting for a CPU; with r above the CPU count every second, processes are queueing for CPU time. b is the number blocked waiting for I/O. si and so show swapping in and out, and anything above zero there for long means memory is short.

The rest can wait until you need them: the memory columns are in KiB, bi and bo count blocks read from and written to disk, in and cs interrupts and context switches per second, and the last columns are the same CPU split as in top. vmstat is a good first command on a machine you do not know, because it shows CPU, memory, swap and disk activity side by side as they change.

The kernel can also say how much time work spent waiting. Pressure stall information (PSI) lives in /proc/pressure, with a file each for cpu, memory and io:

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/pressure/cpu
some avg10=50.40 avg60=34.10 avg300=14.02 total=180777278 full avg10=0.00 avg60=0.00 avg300=0.00 total=0

some avg10=50.40 means that during the last 10 seconds, runnable processes were kept waiting for a CPU about half of the time; avg60 and avg300 cover one and five minutes, and total is the accumulated waiting time in microseconds. For the CPU, the full line (every process stalled at once) is not defined for the whole system and stays at zero. PSI answers "is anything being slowed down?" directly, where the load average only counts processes. It is available on Ubuntu 26.04; the RHEL 10 kernel is built with PSI switched off unless the machine boots with psi=1:

deploy@rocky10 · Rocky Linux 10.2
$ cat /proc/pressure/cpu
cat: /proc/pressure/cpu: No such file or directory

Once the lab stopped the load, the pressure figures fell within seconds while the load average took longer:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop ess-cmd-monitor-load
$ uptime cat /proc/pressure/cpu
11:13:26 up 49 min, 2 users, load average: 1.70, 1.19, 0.87 some avg10=7.31 avg60=24.83 avg300=13.19 total=180965176 full avg10=0.00 avg60=0.00 avg300=0.00 total=0

Twenty-five seconds later avg10 had dropped from 50 to 7, while the one-minute load average had only fallen from 2.23 to 1.70, because it forgets old values gradually. Use PSI or vmstat to see what is happening now and the load average to see the trend.

Memory: read available, not free

deploy@web01 · Ubuntu 26.04 LTS
$ free -h
total used free shared buff/cache available Mem: 3.8Gi 540Mi 765Mi 1.4Mi 2.7Gi 3.3Gi Swap: 0B 0B 0B

free -h shows memory in readable units. total is the RAM the kernel can use. free is memory nothing is using at all, and on a healthy server it is often small, because the kernel keeps recently read file data in RAM (buff/cache) instead of leaving memory empty; most of that cache is handed back as soon as a program needs the memory. available is the kernel's estimate of how much a new program could get without swapping, and it is the number to read: here 3.3 GiB of 3.8. In this version of free, used is simply total minus available. shared is mostly files in memory-backed filesystems such as /tmp. The Swap line is all zeros because cloud images such as this one come without swap. All of these come from /proc/meminfo, in KiB:

deploy@web01 · Ubuntu 26.04 LTS
$ grep -E '^(MemTotal|MemFree|MemAvailable|Cached|SwapTotal):' /proc/meminfo
MemTotal: 3989348 kB MemFree: 783128 kB MemAvailable: 3435200 kB Cached: 2391184 kB SwapTotal: 0 kB

When available approaches zero, the kernel's out-of-memory killer ends a process to free memory, and a service can be killed the same way for exceeding its own memory limit (MemoryMax=) while the machine as a whole still has memory left. Either way the kernel logs the kill, so journalctl -k is where to look when a process vanished without a trace in its own log.

Who is logged in, and who was

On Ubuntu 26.04, who prints nothing, even with two users logged in:

deploy@web01 · Ubuntu 26.04 LTS
$ who echo "who exit status: $?"
who exit status: 0
$ ls -l /run/utmp /var/log/wtmp
ls: cannot access '/run/utmp': No such file or directory -rw-rw-r-- 1 root utmp 2000 Sep 27 11:12 /var/log/wtmp

The default who reads /run/utmp, a file of current logins that dates from long before systemd, and Ubuntu 26.04 no longer creates it: its record format stores times in 32 bits, which overflow in 2038. Its companion history file, /var/log/wtmp, still exists, and sshd still appends logins that have a terminal to it, but Ubuntu no longer installs the tools that read it (below). The current sessions are known to systemd-logind, the service that tracks logins. The GNU who, installed as gnuwho (the first lesson explained the gnu* commands), asks it and does list the sessions:

deploy@web01 · Ubuntu 26.04 LTS
$ gnuwho
ess-cmd-monitor-alice sshd pts/0 2026-09-27 11:12 (127.0.0.1) lima sshd 2026-09-27 10:23

Two standard commands ask logind as well, with more detail:

deploy@web01 · Ubuntu 26.04 LTS
$ w
11:13:26 up 49 min, 2 users, load average: 1.70, 1.19, 0.87 USER TTY FROM LOGIN@ IDLE JCPU PCPU WHAT ess-cmd- pts/0 127.0.0.1 11:12 1:25 0.01s 0.01s -bash lima - 10:23 0.00s 0.12s sshd-session: lima [priv]
$ loginctl list-sessions
SESSION UID USER SEAT LEADER CLASS TTY IDLE SINCE 1 502 lima - 1012 manager-early - no - 108 4181 ess-cmd-monitor-alice - 337410 user - no - 109 4181 ess-cmd-monitor-alice - 337420 manager - no - 4 502 lima - 1505 user - no - 4 sessions listed.

w shows each user, the terminal (pts/0 is an SSH login with a terminal), where the login came from, when it started, how long the user has been idle, and what they are running. It cuts user names to eight characters, so alice appears as ess-cmd-. The lima line is the lab virtual machine tool's own connection. loginctl list-sessions lists the sessions with full user names; the manager sessions belong to each user's own service manager, not to extra logins.

Login history used to come from last (successful logins) and lastb (failed ones), which read /var/log/wtmp and /var/log/btmp. Ubuntu 26.04 no longer installs them:

deploy@web01 · Ubuntu 26.04 LTS (default install)
$ last
-bash: line 1: last: command not found
$ lastlog
-bash: line 1: lastlog: command not found

Every SSH login is still recorded by the SSH server in the journal and in /var/log/auth.log, which the lesson "Logs: the journal and /var/log" covers in detail. To see who logged in, search for the Accepted lines:

deploy@web01 · Ubuntu 26.04 LTS
$ journalctl -u ssh --since 11:12:00 -g Accepted --no-hostname
Sep 27 11:12:00 sshd-session[337410]: Accepted publickey for ess-cmd-monitor-alice from 127.0.0.1 port 43686 ssh2: ED25519 SHA256:4hkbOIx5CnLdunb7cCcdCyc56rCkQDmIXRYxfLh1ids

Each line gives the account, the address it came from and the key that was used. If you want the last command back, the wtmpdb package in Ubuntu's universe repository provides it (as a link to wtmpdb), with a PAM module that records logins in a database that has no 2038 problem. RHEL 10 still ships who, last and lastb:

deploy@rocky10 · Rocky Linux 10.2
$ who
ess-cmd-monitor-alice pts/0 2026-09-27 10:56 (127.0.0.1)
$ last -n 3
ess-cmd- pts/0 127.0.0.1 Sun Sep 27 10:56 still logged in …

Is the clock right?

A wrong clock breaks more than timestamps. TLS certificates are valid only between two dates, one-time login codes depend on the time, and logs from several servers can only be lined up into one timeline if their clocks agree.

deploy@web01 · Ubuntu 26.04 LTS
$ date date -u
Sun Sep 27 11:13:26 UTC 2026 Sun Sep 27 11:13:26 UTC 2026
$ timedatectl
Local time: Sun 2026-09-27 11:13:26 UTC Universal time: Sun 2026-09-27 11:13:26 UTC RTC time: Sun 2026-09-27 11:13:27 Time zone: Etc/UTC (UTC, +0000) System clock synchronized: yes NTP service: active RTC in local TZ: no

date prints local time and date -u UTC; this server's time zone is UTC, as on most servers, so they match. timedatectl adds the hardware clock (RTC), the time zone, and the two lines that matter: NTP service: active means a time synchronisation service is running, and System clock synchronized: yes means it has set the clock from its time servers. On Ubuntu 26.04 that service is chrony, and chronyc sources shows which servers it uses:

deploy@web01 · Ubuntu 26.04 LTS
$ chronyc sources
MS Name/IP address Stratum Poll Reach LastRx Last sample =============================================================================== ^* 185.125.190.122 2 6 377 65 +18ms[ +19ms] +/- 106ms ^+ 185.125.190.123 2 6 377 63 +18ms[ +18ms] +/- 98ms ^+ 91.189.91.112 2 6 377 65 -10ms[-9658us] +/- 195ms ^+ 91.189.91.113 2 6 377 2 -812us[ -812us] +/- 193ms ^- 91.189.91.111 2 6 377 1 +12ms[ +12ms] +/- 208ms

Each line is a time server, and ^ marks a server. The second column is chrony's verdict: * is the source the clock is synchronised to, + a source combined with it, - a usable source that is not currently selected, and ? one it cannot use. Reach is an octal record of the last eight polls: 377 means all eight were answered and 0 means none was. Here every server shows 377. When the clock drifts or timedatectl says no, this is the check to make: if every line shows ^? with Reach 0, no server answers, and the fault lies between the server and its time sources.

Ubuntu 26.04's chrony also uses NTS (Network Time Security, authenticated NTP), which needs outbound TCP port 4460 as well as UDP port 123; the Linux hardening course covers it in "Logs and clocks you can trust".

The first two minutes on an unfamiliar server
How busy
uptime, nproc
load compared with the CPU count
vmstat 1 5
r above CPUs means a queue; watch wa
/proc/pressure/cpu
share of time work waited (Ubuntu)
Memory
free -h
read available, not free
journalctl -k
out-of-memory kills are logged here
People and time
w, loginctl
who is logged in now
journalctl -u ssh -g Accepted
who logged in, from where
timedatectl, chronyc
is the clock synchronised?

Try this

On your practice machine, start three busy loops in the background with yes > /dev/null & three times (yes prints a line endlessly, and the output is thrown away). Wait about fifteen seconds, then run uptime, vmstat 1 3 and cat /proc/pressure/cpu. On a two-CPU machine r shows three or more, id drops to 0 and some avg10 climbs well above zero. Most of the CPU time appears under sy rather than us, because yes spends it inside the kernel writing to /dev/null. Stop the loops with kill %1 %2 %3, confirm with pgrep -c -x yes (it prints 0), and watch avg10 fall over the next half minute while the load average takes several minutes.

Takeaway

Compare the load with the CPU count, read available for memory, and check wa and PSI before blaming the CPU. On Ubuntu 26.04, ask w or loginctl who is logged in and the journal who has been, and make sure timedatectl says the clock is synchronised before trusting any timeline.

Quick check
01A two-CPU server shows a load average of 3.9, 1.6, 1.0. top shows id 0.0 and wa 0.0. What is the most likely situation?
Incorrect — Waiting for I/O would show as wa and processes in state D. Here wa is 0 and the CPUs are fully busy.
Correct — Load above the CPU count with no idle time means a CPU queue, and a one-minute figure far above the fifteen-minute one means it is new.
Incorrect — Load has to be compared with the CPU count. 3.9 on two CPUs means work is waiting.
Incorrect — This server has no swap, and nothing in these numbers points at memory. Check available in free before concluding that.
02On an Ubuntu 26.04 server you are logged in over SSH, yet who prints nothing and exits with status 0. Why?
Incorrect — who used to list SSH logins as well. The difference is where it reads the data from.
Incorrect — who needs no special permission, and it would still show your own session.
Incorrect — No sshd option hides logins from who. The file who depends on is simply not written.
Correct — The old utmp format has a 2038 limit, and current sessions are tracked by systemd-logind.
03timedatectl says "System clock synchronized: no" and "NTP service: active". chronyc sources lists five servers, all marked ^? with Reach 0. What should you check first?
Correct — Reach 0 means none of the last eight polls got a reply, so the problem is between chrony and the servers.
Incorrect — The time zone only changes how times are displayed. Synchronisation is about the clock itself.
Incorrect — chronyc answered with a list of sources, so chronyd is installed and running.
Incorrect — A bad RTC affects the time at boot. Here the service runs but gets no answers.

Related