Test yourself
Advanced Linux internals & tooling
Final exam · 52 questions · answers explained as you pick
Method and processes
10 questions
01On a 2-CPU Ubuntu 26.04 server under load, vmstat 1 5 prints a first row whose us, sy and id columns read 11, 5 and 79, then four rows with us at 99 or 100, id at 0 and r at 4. A colleague's ticket says the CPUs are 79% idle. What went wrong?
Incorrect — The later rows are the per-second samples. The first row is the one that covers a different period: the whole time since boot.
Correct — On a host that was mostly idle since boot, the since-boot row reads mostly idle. The interval rows, us 100 and id 0 with r at 4 on two CPUs, show saturation now.
Incorrect — vmstat's cpu columns are percentages across all CPUs; mpstat -P ALL is the per-CPU view.
Incorrect — vmstat reads the system-wide counters in /proc/stat for every line; nothing is sampled per CPU.
02You install sysstat on a RHEL 10 host on Monday so that sar can show when the nightly slowdowns start. On Tuesday sar -q has no samples from Monday night, and systemctl is-active sysstat-collect.timer prints inactive. Why?
Incorrect — /var/log/sysstat is Ubuntu's path; RHEL uses /var/log/sa/saDD. The path is not the problem: the timer never ran.
Incorrect — sysstat-collect.timer takes a sample every 10 minutes once it runs. The problem is that it is not running.
Incorrect — sar reads today's file by default on both platforms; -f selects the file of another day.
Correct — Collection would begin at the next boot. sudo systemctl enable --now sysstat starts it at once, and is-active then prints active.
03cat /proc/loadavg prints "2.89 2.68 1.34 3/161 245823". What do the last two fields tell you?
Correct — The fraction is runnable tasks over all tasks, and the last field is the most recent PID the kernel allocated.
Incorrect — The fraction counts tasks, not CPUs, and context switches are the ctxt line in /proc/stat.
Incorrect — Blocked tasks are procs_blocked in /proc/stat, and tasks created since boot are its processes line.
Incorrect — Zombies are not counted in this file, and uptime lives in /proc/uptime.
04On a cloud VM, lsblk -t shows ROTA 1 for vda and /sys/block/vda/queue/rotational reads 1, although the provider sells the volume as SSD storage. A tuning script uses that flag to decide how fast the disk is. What should you make of the 1?
Incorrect — The kernel measures nothing here; the value is whatever the driver declares.
Incorrect — RAID does not set this flag, and nothing in the output shows a RAID layer.
Correct — The lab's virtio disk reports 1 while it is a file on the host's SSD. Measure latency with iostat or fio instead of trusting the flag.
Incorrect — The scheduler has its own file, queue/scheduler; rotational comes from the driver.
05A service keeps running after kill -TERM. Its /proc/PID/status shows SigBlk 0000000000000200, SigIgn 0000000001001000 and SigCgt 0000000000004002. What happened to the SIGTERM?
Incorrect — Blocked signals are in SigBlk, and 0x200 is bit 9, signal 10 (SIGUSR1), not SIGTERM.
Correct — Bit n-1 stands for signal n, so 0x4002 is SIGINT and SIGTERM. The handler ran and chose to carry on; any line it logs confirms it.
Incorrect — SigIgn 0x1001000 is bits 12 and 24, signals 13 and 25 (SIGPIPE and SIGXFSZ), not signal 15.
Incorrect — Standard signals merge duplicates of a pending signal; a single SIGTERM is not lost, and the mask shows a handler.
06Two dd readers hang on a storage path that stopped answering. ps shows both in D, one with wchan submit_bio_wait (it reads with iflag=direct) and one with folio_wait_bit_common (it reads through the page cache). After sudo kill -KILL on both, only the direct reader is left. Why did the other one end?
Correct — Both show as D, but TASK_KILLABLE ends on SIGKILL, while TASK_UNINTERRUPTIBLE waits for the I/O to complete with SIGKILL left pending in ShdPnd.
Incorrect — The device stayed suspended the whole time; the buffered reader left because of the SIGKILL, not because its read completed.
Incorrect — A signal is delivered at once, and the direct reader's status shows it pending. The difference is the kind of sleep.
Incorrect — Direct I/O readers receive signals like any task; the delay comes from the plain uninterruptible wait.
07On Ubuntu Server 26.04 you pin a web scope and a batch scope, one CPU-bound worker each, to CPU 1 and renice the batch worker to 19. pidstat still shows 50.20% and 50.00%. The same test on Rocky Linux 10.2 with nice 10 splits about 90 to 10. What explains the Ubuntu result?
Incorrect — Nice works for every fair task: two workers inside one Ubuntu scope split 90.24 to 9.76 at nice 0 and 10.
Incorrect — Nice still sets the weight under EEVDF (se.load.weight 112640 at nice 10); a shorter slice changes latency, not the share.
Incorrect — Raising your own nice value is allowed, and renice printed the new priority 19. Lowering it again is what needs privilege.
Correct — multipathd.service sets CPUWeight=1000, so Ubuntu's system.slice passes cpu down and nice ranks tasks within one group. Rocky's system.slice passes memory and pids alone.
08A tuning guide for low-latency services tells you to put kernel.sched_latency_ns and kernel.sched_min_granularity_ns in a file under /etc/sysctl.d. On Ubuntu 26.04, sysctl answers "cannot stat /proc/sys/kernel/sched_latency_ns: No such file or directory". What is the situation?
Incorrect — sched_ext is a separate scheduling class for BPF schedulers; loading one does not create these sysctls.
Incorrect — sysctl files set entries that exist; they cannot create one the kernel no longer has.
Correct — EEVDF has a base slice instead (1.4 ms on this 2-CPU VM), readable in /sys/kernel/debug/sched, a debugging interface rather than a supported setting.
Incorrect — A permission problem is reported as permission denied; "cannot stat" means the file does not exist.
09Two unpinned CPU-bound workers run on a 2-CPU server. vmstat shows r 2 and us 100, mpstat shows both CPUs near 100%, /proc/pressure/cpu reads some avg10=1.85, and pidstat shows the workers' %wait at 0.60 and 0.00. Is the CPU a bottleneck to act on?
Correct — Each worker has a CPU to itself, so pressure and %wait stay low. The load average settles at 2, the CPU count, on a healthy machine.
Incorrect — Busy CPUs are utilisation. Saturation is waiting, and the pressure and %wait figures show very little of it.
Incorrect — USE defines saturation as work that has to queue. r equal to nproc means each runnable task has a CPU.
Incorrect — No limit is set on these workers. PSI measures stall time, and a low value means little stall.
10Two workers pinned to CPU 1 run in perf-cpu-pinned.scope. Its cpu.pressure reads "some avg10=93.83" and "full avg10=0.00 avg60=0.00 avg300=0.00 total=12773". Why is full near zero while some is so high?
Incorrect — At system level CPU full is reported as 0, but cgroup files do count it, which is why total here is 12773 and not 0.
Incorrect — Throttling is one cause of stalls, not a condition for counting full.
Incorrect — The three averages use the same data over different windows, and avg300 of full is also 0.00.
Correct — some means at least one task in the group waited; full means all of them waited, which two workers sharing one CPU seldom do.
10 questions · explanations appear as you answer
Boot and systemd
7 questions
01An Ubuntu 26.04 VM boots more slowly than you would like. systemd-analyze blame is led by sys-devices-virtual-misc-rfkill.device at 1.678s, dev-vport1p0.device at 1.653s and dev-disk-by\x2dlabel-BOOT.device at 1.488s. critical-chain shows graphical.target @2.489s behind snapd.seeded.service @1.946s +541ms. Which unit do you look at first?
Incorrect — blame ranks every unit started since boot by its own time. None of these device units is on the chain the default target waited for.
Incorrect — Their times overlap, and nothing on the critical chain waits for them, so adding them up measures nothing about the boot.
Correct — critical-chain shows the path that gated the target; units off that path can take any time without delaying the boot.
Incorrect — The missing firmware and loader figures do not affect the userspace timings; they are simply absent.
02A server rebooted overnight and nobody admits to it. journalctl -b -1 -n 3 ends with "systemd-journald[1123911]: Received SIGTERM from PID 1 (systemd-shutdow)." and "systemd-journald[1123911]: Journal stopped". What does that tell you?
Incorrect — A panic halts the system without running systemd-shutdown; the journal would end mid-stream.
Correct — journald was stopped by systemd-shutdown in the normal order. Next, find who or what asked for the reboot: a user, an update tool or a timer.
Incorrect — -b -1 selects the previous boot's entries; a journald restart would not end that boot's log.
Incorrect — A power cut gives systemd no chance to run shutdown; the log would simply stop.
03On an Ubuntu 26.04 cloud image you add loglevel=4 to GRUB_CMDLINE_LINUX_DEFAULT in /etc/default/grub and run update-grub, but the linux line in /boot/grub/grub.cfg still ends in "console=tty1 console=ttyAMA0" without it. Why?
Incorrect — The ESP grub.cfg is a pointer to /boot/grub/grub.cfg, which is the file update-grub regenerates.
Incorrect — update-grub writes the file at once; the reboot is what makes the kernel use it.
Incorrect — BLS entries managed by grubby are how RHEL's GRUB works; Ubuntu builds grub.cfg from /etc/default/grub.
Correct — update-grub reads /etc/default/grub, then each file in /etc/default/grub.d/, and the last assignment wins. Put your change in a drop-in that sorts later.
04app.service has Requires=db.service and nothing else in [Unit]. db.service is Type=notify and needs five seconds to become ready. Right after sudo systemctl start app, systemctl is-active db app prints "activating" and "failed", and app logged "database not ready". What is missing?
Correct — Requires= pulls db into the transaction but orders nothing, so both started together and app lost the race. With After= the start took five seconds and app connected.
Incorrect — Wants= and Requires= differ in how a failure is treated; neither of them orders the two units.
Incorrect — That would make other units wait for app; it does nothing to make app wait for db.
Incorrect — static only means the unit cannot be enabled. It is still started as a requirement, at the same moment as app.
05While a start job for sd-units-app waits behind its database, sudo systemctl stop --job-mode=fail sd-units-app prints "Transaction for sd-units-app.service/stop is destructive (sd-units-app.service has 'start' job queued, but 'stop' is included in transaction)". What did systemd do?
Incorrect — With --job-mode=fail nothing was queued or cancelled; the whole stop transaction was refused.
Incorrect — That describes the default mode, replace. With fail, systemd refuses to cancel a queued job.
Correct — fail rejects any transaction that conflicts with queued jobs; without the option, replace would have swapped the start for the stop.
Incorrect — A loop is reported as an ordering cycle, and systemd never edits unit files.
06You fix a mistyped option on the /boot line of /etc/fstab. systemctl cat boot.mount still shows the old unit, headed "# /run/systemd/generator/boot.mount" and "# Automatically generated by systemd-fstab-generator". What has to happen before systemd sees your change?
Incorrect — The generator rewrites that directory at each reload and boot, so a manual edit there is lost.
Incorrect — Generators also run at every daemon-reload, which is what the command is for.
Incorrect — A copy there would override fstab for good, and later fstab edits would be ignored.
Correct — The unit is generated from fstab (SourcePath=/etc/fstab), so it changes only when the generator runs again.
07Ubuntu 26.04 listens for SSH through ssh.socket (Accept=no), and the listener's Send-Q reads 4096. An OpenSSH update restarts ssh.service, and a client connects during the second in which no sshd process exists. What happens to that connection?
Incorrect — PID 1 still holds the listening socket, so the port stays open while the daemon restarts.
Correct — systemd created the socket and keeps it across service restarts; the kernel completes the handshake and queues the connection, up to the backlog of 4096 that Send-Q shows.
Incorrect — That is Accept=yes. ssh.socket has Accept=no: one ssh.service instance receives the listening socket and accepts every connection.
Incorrect — The socket never stopped listening, so the kernel answers the first SYN; nothing has to be retransmitted.
7 questions · explanations appear as you answer
Memory and resource control
8 questions
01journalctl -k shows "k-mem-hog invoked oom-killer" and then "oom-kill:constraint=CONSTRAINT_MEMCG,...,oom_memcg=/system.slice/k-mem-oom.service,task_memcg=/system.slice/k-mem-oom.service,task=k-mem-hog". free on the host shows gigabytes available. What does the report tell you?
Correct — CONSTRAINT_MEMCG means a cgroup limit was hit, and oom_memcg names that cgroup. A machine-wide OOM says CONSTRAINT_NONE.
Incorrect — free counts the host's memory correctly; the kill happened inside one cgroup, not at host level.
Incorrect — The first line names the task whose allocation failed. The victim is chosen by badness and can be another task.
Incorrect — systemd-oomd is not installed on Ubuntu Server 26.04, and its kills are not kernel oom-kill lines.
02After the page cache is dropped on a lab machine, a program that touches each page of a 64 MiB file reports "pass 1: 435 minor, 1 major". With readahead turned off (MADV_RANDOM) the same pass reports "pass 1: 0 minor, 16384 major" and then "pass 2: 1024 minor, 0 major". What does the difference show?
Incorrect — Minor faults are handled by the kernel too; they simply need no I/O.
Incorrect — Each of the 16,384 pages faulted once; the count matches the file size in pages.
Correct — Sequential access let the kernel read ahead, so later faults found their pages cached. Without it each page is a separate major fault, which is how a database whose working set misses the cache becomes disk-bound.
Incorrect — The second pass shows 1024 minor and 0 major faults: the pages were cached and served from memory.
03On an Ubuntu 26.04 host /proc/meminfo shows CommitLimit 1994676 kB and Committed_AS 2587248 kB, and a program that maps 3 GiB of private memory on this 3.8 GiB machine still succeeds. Why is nothing refused?
Incorrect — The limit has nothing to do with who runs the program; it depends on the overcommit mode.
Correct — vm.overcommit_memory=0 is a heuristic: an 8 GiB request, more than RAM and swap together, was refused with ENOMEM. Shortages show up later, when pages are touched.
Incorrect — Committed_AS totals private writable mappings, which is exactly what the program added.
Incorrect — This machine has no swap; reserving address space costs neither RAM nor swap until pages are touched.
04A batch slice with CPUQuota=50% runs slowly. In its cgroup directory cpu.max reads "50000 100000", and cpu.stat shows nr_periods 63, nr_throttled 63 and throttled_usec 9225117. The host's two CPUs are mostly idle. What limits this workload?
Incorrect — Run-queue waiting is %wait and schedstat. Throttled time is time the group was held back by its quota.
Incorrect — nr_throttled cannot exceed nr_periods; equal values mean the quota ran out in every period.
Incorrect — cpu.stat belongs to the CPU controller; memory throttling appears as high in memory.events.
Correct — 50 ms per 100 ms period is used up every time. Lifting the quota with set-property CPUQuota= let each worker use a full CPU.
05A scope started with -p IOReadBandwidthMax="/var/tmp 8M" reads a file at 7995 kB/s, and its io.max reads "253:0 rbps=8000000 wbps=max riops=max wiops=max". Nothing else uses the disk. What does this show?
Correct — /var/tmp resolved to vda (253:0), 8M means 8,000,000 bytes a second, and io.max is not work-conserving: the read took 8.4 s on an idle disk.
Incorrect — io.max limits a device; the path only told systemd which device. Everything the scope reads from vda is capped.
Incorrect — The disk uses mq-deadline, and io.max works with any scheduler. It is weights that need BFQ or iocost.
Incorrect — rbps is an absolute ceiling, and no other cgroup was using the disk.
06A shell started with sudo unshare --pid --fork prints "my PID: 1" for itself, yet ps -e --no-headers | wc -l inside it prints 125, every process on the host. Why does ps see them all?
Incorrect — The shell is already PID 1 in the new namespace; the namespace is in effect.
Correct — With --mount-proc, unshare also creates a mount namespace and mounts a new /proc, and ps then sees itself as PID 1.
Incorrect — What ps sees follows the procfs mount, not the user; root inside with a fresh /proc sees the namespace alone.
Incorrect — The caller stays where it is; its next child enters the new namespace, which is why --fork is needed.
07On Rocky Linux 10.2, as an ordinary user, unshare --user --map-root-user shows uid=0(root), a uid_map of "0 1001 1" and CapEff 000001ffffffffff. Yet cat /etc/shadow inside fails with Permission denied, and ls -ln shows the file owned by 65534. Why?
Incorrect — The owner appears as 65534 because of the ID mapping, and the capability check fails before SELinux has to decide anything.
Incorrect — CapEff is already full in this process, as the status line shows; no second exec is needed.
Correct — The namespace maps UID 1001 alone, so the owner of /etc/shadow shows as the overflow ID 65534 and the capabilities do not apply to it.
Incorrect — Real root reads it through CAP_DAC_OVERRIDE; what differs is which user namespace owns the capability.
08You delegate a unit's cgroup and write "+memory +pids" to its init/cgroup.subtree_control, where the unit's processes live. The write fails with "Device or resource busy (os error 16)", while the same write to the unit's own, empty cgroup succeeds. Why?
Incorrect — Writing a controller that is already on is accepted; the refusal comes from the processes in init.
Incorrect — Delegation hands over the whole subtree; the lab created and configured box beneath the unit.
Incorrect — The Rust tee only prints the error number its own way; the kernel returned EBUSY to the write.
Correct — That is why DelegateSubgroup=init put the processes in a child and left the unit's cgroup empty for subtree_control.
8 questions · explanations appear as you answer
Files and storage
10 questions
01A shell opens /etc/passwd on fd 3, duplicates it with exec 4<&3, opens it again on fd 5, and reads one line through fd 3. fdinfo then shows pos 32 for fd 3, pos 32 for fd 4 and pos 0 for fd 5. Why did fd 4 move?
Incorrect — fd 5 refers to the same inode and stayed at 0, so the offset is not kept per inode.
Incorrect — The page cache holds file data, not positions; readahead does not move anyone's offset.
Incorrect — Bash does nothing to fd 4; the kernel reports an offset that the two descriptors share.
Correct — open() creates a description with its own offset (fd 5); dup() and fork() share one, which is also why a child's reads move its parent's position.
02A log shipper on a 4 GiB server stops picking up new directories, and its log shows ENOSPC (No space left on device) from inotify_add_watch. df -h shows the disks under 60% full, df -i shows plenty of free inodes, and /proc/sys/fs/inotify/max_user_watches reads 30890. What is the likely cause?
Correct — The watch limit, sized to about 1% of memory, fails inotify_add_watch with ENOSPC; the instance limit fails inotify_init with EMFILE instead.
Incorrect — A full journal forces a checkpoint that writes blocks back; it does not fail an inotify call.
Incorrect — A cgroup memory limit leads to reclaim or an OOM kill, not to an ENOSPC error.
Incorrect — df does count blocks of deleted-but-open files, which is why it disagrees with du; here it shows free space.
03Straight after echo hello > /mnt/st-fs-ext4/small, filefrag -v shows its extent at physical_offset 0 with the flags "last,unknown_loc,delalloc,eof". After sync it shows block 4171 and the flags "last,eof". What did the first output mean?
Incorrect — The write succeeded: the file has a size of 6 bytes and its data in the page cache.
Incorrect — A hole has no extent at all; delalloc marks reserved space that has no location yet.
Correct — ext4 and XFS pick physical blocks when the data is written back and they know how much there is, which keeps files contiguous.
Incorrect — Inline data is a separate feature that is not in use here; delalloc means the block is reserved but not placed.
04After a crash, a colleague runs sudo fsck /dev/sdb1 on an unmounted XFS data volume. It prints a two-line pointer ending "see xfs_repair(8)." and exits 0, and they mark the volume as checked. What is wrong?
Incorrect — On XFS, fsck.xfs does not check anything, so its 0 says nothing about the filesystem.
Correct — The lab's zeroed directory block passed fsck with status 0 and was found by xfs_repair -n. If the log is dirty, mount and unmount first.
Incorrect — -y is an e2fsck answer mode; fsck.xfs ignores it and prints the same pointer.
Incorrect — Mounting replays the log, but checking and repairing need the filesystem unmounted.
05A deploy script on Ubuntu 26.04 runs "sync /srv/app/state.db" after each update, meant to flush that one file. strace -e trace=sync,syncfs,fsync,fdatasync shows the command calling sync(), not fsync(). What does that mean?
Incorrect — On Linux sync() waits for the writes, and it flushes this file along with every other filesystem.
Incorrect — That is GNU sync, installed as gnusync; Ubuntu 26.04's default sync is the Rust uutils version, which calls sync().
Incorrect — There is no such fallback; which call is made depends on which sync implementation runs.
Correct — The result is durable but heavier than intended. gnusync calls fsync() on the file, and the Rust sync --data calls fdatasync().
06A mounted ext4 volume is 96% full. sudo lvextend --size +200M stlvm/data reports "Size of logical volume stlvm/data changed from 200.00 MiB (50 extents) to 400.00 MiB (100 extents)", but df -h still shows Size 172M. What is the next step?
Incorrect — The filesystem holds the old size, not the kernel; mounting again does not change it.
Incorrect — The VG had free extents, which is how lvextend succeeded; the LV is already 400 MiB.
Correct — The filesystem was made for 200 MiB and does not know the device grew; resize2fs took df to 359M. lvextend --resizefs does both steps.
Incorrect — lvextend updates metadata and the device-mapper table; nothing is copied.
07After a volume was grown, sudo dmsetup table stlvm-data shows "0 516096 linear 7:0 2048" and "516096 303104 linear 7:1 2048". A colleague says the volume is now mirrored across both disks. What does the table actually say?
Correct — Two linear segments concatenate the disks: spanning adds capacity, not redundancy. lvs --segments shows the same layout.
Incorrect — Mirroring uses a raid or mirror target; linear sends each sector range to exactly one device.
Incorrect — A snapshot appears as snapshot and snapshot-origin targets; these are two linear segments of one LV.
Incorrect — A striped LV uses the striped target; linear segments are simply placed one after the other.
08A test disk that serves four requests at a time shows, with four readers, 360.60 r/s, r_await 10.99 ms and aqu-sz 3.96. With eight readers iostat shows 348.10 r/s, r_await 22.85 ms, aqu-sz 7.95 and %util 99.84. What changed?
Incorrect — %util was 99.94 with four readers while latency stayed at the service time: full use without queueing.
Correct — Four requests are served and four wait for a slot, each for about one service time. Latency above the quiet-time value with flat throughput is the signature.
Incorrect — The device still takes 10 ms per request; the extra time is spent waiting in the queue for a slot.
Incorrect — The readers wait on the disk, as I/O pressure and iodelay show; the CPUs are not busy.
09A VM has a slow test disk (no request finishes in under 10 ms) and ordinary system I/O on its boot disk. biolatency.bt shows a few dozen requests between 64 µs and 4 ms and a peak of 3273 requests in [16K, 32K), and fio reports about 22.8 ms per read. A colleague reads the fast requests as proof that the test disk is often quick. How do you settle it?
Incorrect — The units convert directly: 16K to 32K µs is 16 to 32 ms, in line with fio. The question is which device each request went to.
Incorrect — More time adds counts to both populations; it cannot tell you which device each one belongs to.
Correct — Keyed on args.dev, [251, 0] (nullb0) had every request between 8 and 64 ms, and [253, 0] (vda) held the twenty fast ones. biolatency.bt alone keys on the sector and cannot say which device a request went to.
Incorrect — With none the readers still wait, before the block layer starts timing; fio still reports about 22.4 ms per read.
10df -h /mnt/perf-io shows 41M used and du -sh shows 20K. lsof +L1 lists /mnt/perf-io/app.log (deleted), NLINK 0, open on fd 1w by two processes, sh and sleep. Why do two processes hold the same deleted file?
Correct — A child inherits its parent's descriptors, so the blocks are freed when both close it; stopping the writer's unit ends both.
Incorrect — Neither opened it again: sh redirected its output once with >> and sleep inherited that descriptor.
Incorrect — lsof lists processes that hold the file open, and both of these do.
Incorrect — NLINK 0 says no names remain; a hard link would keep a name, and du would count it.
10 questions · explanations appear as you answer
Networking
7 questions
01The first line of /proc/net/softnet_stat on a 2-CPU host begins "00035a27 00000000 0000000e". net.core.netdev_budget is 300 and netdev_budget_usecs 2000. What does the line say about CPU 0?
Incorrect — Backlog drops are the second column, which is 0 here.
Correct — The columns are processed, dropped and time_squeeze, in hexadecimal. A few squeezes are normal; a count that climbs with traffic, or any drops, means packet processing cannot keep up.
Incorrect — The first column counts packets processed since boot; it is not a queue length.
Incorrect — Checksum and framing failures are errors in ip -s link, not softnet_stat columns.
02On an Ubuntu 26.04 cloud VM, sysctl net.core.default_qdisc prints "net.core.default_qdisc = fq_codel", but tc -s qdisc show dev eth0 prints "qdisc pfifo_fast 0: root refcnt 2 bands 3". Why does eth0 not use fq_codel?
Incorrect — default_qdisc has no speed condition; the Rocky VM's interface got fq_codel from the same setting.
Incorrect — A root qdisc and default_qdisc both belong to the transmit (egress) path; received packets do not pass a root qdisc.
Incorrect — The setting does not look at the kind of device; the value is taken when the interface is created or comes up.
Correct — A qdisc is chosen when an interface comes up, so an interface raised early keeps the built-in pfifo_fast for as long as it stays up.
03On Rocky Linux 10.2, getent hosts _gateway prints "192.168.5.2 _gateway", yet no DNS server has that name. Where does the answer come from?
Correct — RHEL's hosts line lists files, dns and myhostname, in that order; the module answers for the machine's own name and for _gateway, the current default gateway.
Incorrect — RHEL 10 does not install systemd-resolved; resolv.conf points straight at the upstream server.
Incorrect — NetworkManager writes resolv.conf, not /etc/hosts, and the files source did not supply this name.
Incorrect — A DNS server has no name for your gateway; a dig for it would return NXDOMAIN.
04curl from a client to http://198.51.100.10:8080/ fails after 0 ms with "Could not connect to server". A capture on the server's interface shows the client's SYN (Flags [S]) and 198.51.100.10.8080 answering with Flags [R.]. What does the reset tell you?
Incorrect — A drop rule produces silence and repeated SYNs. A reset means something actively refused the connection.
Incorrect — A lost SYN is sent again for seconds; curl failed within a millisecond because an answer arrived.
Correct — Check ss -ltnp on the server (in the lab the service was bound to 127.0.0.1) and look for a rule that rejects with a TCP reset; a load balancer in front would also answer this way.
Incorrect — A SYN carries no payload; MTU problems show up later, with full-size segments.
05After a change on a router, transfers to a server fall from about 97 Gbit/s to 2.3 Mbit/s. mtr -i 0.2 -c 50 shows 10.0% loss and a 43.1 ms average at the server, and iperf3 shows 41 retransmissions with a Cwnd between 7.07 and 11.3 KBytes. How can a few percent of loss cost that much?
Incorrect — The window in iperf3's output is the sender's congestion window; the receiver offered much more (snd_wnd:444928 in ss -ti).
Correct — A window of a few kilobytes per round trip gives roughly the throughput measured. Loss does far more damage than its percentage suggests.
Incorrect — 41 retransmissions out of about 1,000 segments is a small share of what was sent; the window is the limit.
Incorrect — mtr sends a few probes a second, which cannot take up gigabits.
06On the router, tc -s qdisc show dev nd-r1 reports "dropped 72" for its netem qdisc, while ip -s link show nd-r1 shows 0 dropped and 0 missed on both the RX and TX lines. Which reading is right?
Correct — Each counter belongs to one stage. On a real host, loss inside the network is in none of your counters, and TCP retransmissions are the evidence.
Incorrect — Delayed packets are sent, not counted as dropped; the 72 were discarded.
Incorrect — The TX line has its own dropped counter; it stays at 0 because the drop happened above the driver.
Incorrect — The qdisc's counters start when it is created, and both were read while it ran.
07After a day of tests between two namespaces, nstat -asz on the client shows TcpRetransSegs 117 and TcpOutSegs 42225606. A colleague wants to open a ticket about the retransmissions. What is the right reading?
Incorrect — A veth pair has no physical layer to fail; drops in a queue, or an emulated lossy link, explain them.
Incorrect — Corruption would show up as errors. Retransmissions follow drops, and overflowing queues cause them.
Incorrect — 117 segments out of 42 million cannot halve anything; nothing points at congestion control.
Correct — A retransmission count means little until it is compared with TcpOutSegs; nstat without -a gives the change over an interval.
7 questions · explanations appear as you answer
Tracing, profiling and tuning
10 questions
01Jobs dropped into a spool directory sometimes wait up to half a second before a worker picks them up, and the worker uses almost no CPU. sudo strace -T -p on it shows openat, getdents64 and close taking microseconds, then clock_nanosleep(...) = 0 <0.501921>, over and over. What is the diagnosis?
Incorrect — The second getdents64 returns 0 entries, which is how the end of a directory is read, and both calls take microseconds.
Incorrect — The sleep ends on time, about 0.5 s each round; the time is the requested sleep, not a late wake-up.
Correct — Nothing is slow: the program asks to sleep. The fix is in the design, a shorter interval or inotify to be woken when a file arrives.
Incorrect — Tracing adds microseconds per call; it cannot turn into half-second sleeps that the program asks for.
02On Ubuntu 26.04, sudo perf trace -e openat -- cat /etc/hostname prints lines such as "openat(dfd: CWD, filename: 0x9888e680, flags: RDONLY|CLOEXEC, ...) = 3". You need the file names. What do you use?
Correct — Ubuntu's perf is built without those BPF skeletons, so it shows pointer values. strace prints the strings, at a much higher cost per call.
Incorrect — The names are data in the process, not symbols, and sudo-rs on Ubuntu ignores -E anyway.
Incorrect — perf trace records every event from the system call tracepoints; it does not sample.
Incorrect — -s prints counts and times per system call, with no arguments at all.
03Profiling your own program with perf record -F 99 -g -- ./app stops with "perf_event_paranoid setting is 4" and advice to change /etc/sysctl.conf. This happens on the Ubuntu 26.04 box, while a RHEL 10 box accepts the identical command. How do you proceed on Ubuntu?
Incorrect — Ubuntu 26.04 has no /etc/sysctl.conf, and lowering the setting host-wide for one profile is a policy change the hardening course weighs separately.
Correct — 4 is an Ubuntu addition to the upstream scale, which ends at 2. A one-off sudo leaves the host's policy alone.
Incorrect — perf is installed (linux-perf, version 7.0.14); it ran and printed the error.
Incorrect — adm grants access to logs; no group membership changes perf_event_paranoid.
04A profile recorded with -F 5 captured 21 samples and shows checksum at 95.24% and parse_request at 4.76%. A -F 99 profile of the same program has 468 samples with 91.67% and 8.33%, and another 468-sample run of the same code gives 95.09% and 4.91%. What can you say about parse_request?
Incorrect — The -F 5 figure is one sample out of 21; agreeing with one other run by chance proves nothing.
Incorrect — The same code gave both; a few points of difference between runs on a shared VM is noise until a repeat confirms it.
Incorrect — A sample records where the program was at one instant; spacing them further apart tells you less, not more.
Correct — With 21 samples each one is almost 5% of the profile, too coarse to tell 3% from 8%. Aim for hundreds of samples, by running longer, and repeat before trusting small differences.
05A runbook written on Ubuntu says to run sudo runqlat-bpfcc 5 1. On a RHEL 10 host with bcc-tools installed, the shell finds neither runqlat-bpfcc nor runqlat. Where is the tool?
Incorrect — RHEL ships bcc-tools (131 entries), and the ten tools the lab tried on Rocky all worked.
Incorrect — RHEL uses no suffix at all, and its bcc tools are not in /usr/sbin.
Correct — On RHEL 10 the bcc tools also compile against the 6.12 kernel, unlike runqlat-bpfcc on Ubuntu 26.04's kernel 7.0.
Incorrect — perf sched latency summarises delays per task; it is a separate tool from runqlat's histogram.
06You want to know which files a process makes the kernel sync. bpftrace -lv shows tracepoint:syscalls:sys_exit_write with just __syscall_nr and ret, and fentry:vmlinux:vfs_fsync_range with "struct file * file", start, end and datasync. Which probe answers the question, and at what cost?
Incorrect — A tracepoint exposes the arguments its authors defined, and this one has a number and a return value.
Correct — fentry reads arguments through BTF, so a one-liner can follow the file to its name; kernel function names and arguments are not a stable interface.
Incorrect — The kernel knows which file it syncs; fentry reads it, as the lab's count by file name showed.
Incorrect — BTF describes the types of this kernel; a later kernel can rename or change the function.
07A hung service is killed with SIGABRT and its core opened in gdb. Thread 2 (LWP 257225) waits in pthread_mutex_lock on stats_lock at dt-worker.c:14, thread 1 (LWP 257216) on queue_lock at line 31, and print stats_lock.__data.__owner gives 257216 while queue_lock's owner is 257225. What is the bug, and the fix?
Incorrect — Both owners are live threads in the core, each waiting for the other's lock.
Incorrect — Both use PTHREAD_MUTEX_INITIALIZER, and their owner fields hold real thread IDs.
Incorrect — Each thread waits for a lock the other holds, so no wake-up is due until one of them releases.
Correct — Each thread holds what the other wants. Making flush_queue take stats_lock before queue_lock, as collect_stats does, removes the cycle.
08On RHEL 10 a program segfaults and bash prints "(core dumped)". ulimit -c prints unlimited and core_pattern is "|/usr/lib/systemd/systemd-coredump %P %u %g %s %t %c %h %d %F". How do you get a backtrace?
Correct — systemd-coredump stores the core compressed in /var/lib/systemd/coredump, logs a stack trace in the journal, and coredumpctl list shows it with COREFILE present.
Incorrect — core_pattern is a pipe to systemd-coredump, so no plain core file is written in that directory.
Incorrect — That is Ubuntu's apport directory; RHEL sends cores to systemd-coredump.
Incorrect — systemd-coredump keeps the core of any program whose core limit allows one and whose core fits under ExternalSizeMax (1G on RHEL); this small program qualifies.
09A database vendor asks for transparent huge pages to be set to madvise on your RHEL 10 hosts, which show "[always] madvise never". A colleague adds a line for it to /etc/sysctl.d/60-db.conf. After a reboot /sys/kernel/mm/transparent_hugepage/enabled still shows [always]. Why?
Incorrect — Nothing in sysctl.d sets THP, early or late.
Incorrect — TuneD's own configuration lets /etc/sysctl.d win over a profile's sysctls, and THP is not a sysctl, so the line set nothing in the first place.
Correct — sysctl.d covers /proc/sys. On a RHEL host running TuneD (a standard install does), a custom profile that includes the active one sets it; otherwise transparent_hugepage= on the kernel command line.
Incorrect — THP is adjustable on both; RHEL defaults to [always] and Ubuntu to [madvise].
10You install and enable TuneD on a Rocky Linux 10.2 VM. tuned-adm recommend picks throughput-performance although the host is a VM, sudo virt-what prints nothing, and tuned-adm verify reports "Verification failed", with log lines such as "verify: failed: device cpu0: 'boost' = 'None', expected '1'". What do you conclude?
Incorrect — The other settings did change (swappiness 60 to 10, read-ahead 128 to 4096 KiB); boost was the one that could not apply.
Correct — TuneD detects VMs with virt-what, which does not recognise this hypervisor. Check what an automatic choice was based on, and note that the profile changed several settings at once, none of them measured.
Incorrect — The log names the cause: the virtual CPUs have no boost control, and a reboot does not add one.
Incorrect — TuneD ran and applied most of the profile; one setting could not apply on these virtual CPUs.
10 questions · explanations appear as you answer