Boot: firmware, boot loader, kernel and initramfs
From power-on to systemd, and timing it.
A server that does not come back after a kernel update, or takes minutes to boot, has failed at one stage of a fixed chain: firmware, shim, boot loader, kernel, initramfs, systemd. This lesson follows that chain on running Ubuntu 26.04 and RHEL 10 machines, using the evidence each stage leaves behind: UEFI boot entries, the files on the EFI system partition, the kernel command line, the kernel log, the dracut initramfs and systemd's boot timing. You will change a kernel parameter the supported way on each distribution, find the unit that gates the boot, and know which levers exist when a machine stops before you can log in. Nothing here reboots the lab machines.
Firmware, shim and the boot entry
UEFI firmware keeps its boot entries in non-volatile variables, each naming a file on the EFI system partition (ESP), a small FAT partition that Linux mounts at /boot/efi. /sys/firmware/efi exists only when the running system was started by UEFI firmware, and efibootmgr lists the entries.
BootCurrent: 0002 is the entry this boot used: "Ubuntu", which starts \EFI\ubuntu\shimaa64.efi on partition 15, the ESP. The two "UEFI Misc Device" entries are ones the firmware created for the VM's disks. On an x86_64 server the files are shimx64.efi and grubx64.efi.
shimaa64.efi is the first stage. With Secure Boot on, the firmware only runs code signed by a key in its database, and most x86 hardware ships with Microsoft's UEFI certificates there. shim is signed by Microsoft and carries the distribution's own certificate (Canonical's on Ubuntu, Red Hat's on RHEL), which it uses to verify grubaa64.efi and then the kernel. mmaa64.efi is MokManager, the console tool for enrolling your own Machine Owner Keys, for example to sign a third-party kernel module. \EFI\BOOT\BOOTAA64.EFI is a copy of shim at the path firmware tries when no boot entry works; it then runs fbaa64.efi, which recreates the missing entries from BOOTAA64.CSV.
This VM's firmware does not implement Secure Boot at all, so mokutil says so and the kernel logs it as disabled; on a machine with Secure Boot on, mokutil --sb-state prints SecureBoot enabled. A lab VM cannot show you signature checks, MokManager prompts or TPM measurements (hashes of each boot stage recorded in the Trusted Platform Module, a security chip); they only happen on firmware that supports them. RHEL shows the same. There, sudo bootctl status summarises the firmware, the ESP and the EFI entries in one view; on Ubuntu the command comes in the systemd-boot-tools package, which a default Ubuntu Server install does not include:
command -v prints nothing and dpkg knows no such package. On RHEL, bootctl also warns about the grub_* lines in the boot entries, which it does not understand; those lines are left out below.
GRUB and the kernel command line
The GRUB configuration on the ESP is only a pointer: it finds the /boot filesystem by UUID and loads the real configuration from there.
On Ubuntu, /boot/grub/grub.cfg is generated by update-grub from /etc/default/grub, then every file in /etc/default/grub.d/, and the scripts in /etc/grub.d/; never edit it by hand, because the next kernel update regenerates it. Later files override earlier ones, which is why the cloud image's drop-in replaces quiet splash with two consoles. GRUB_TIMEOUT_STYLE=hidden hides the menu until the timeout passes; pressing Esc or holding Shift during it shows the menu, but with GRUB_TIMEOUT=0, as on this image, there is no such window, so set a timeout before you need the menu.
The linux line in grub.cfg is what GRUB will pass next time. /proc/cmdline is what the running kernel actually received, with BOOT_IMAGE added by GRUB, and it is the only reliable answer to "is this parameter active?" RHEL does it differently: its GRUB reads Boot Loader Specification (BLS) entries, one small file per installed kernel in /boot/loader/entries/, managed with grubby.
The entries directory is readable only by root, so the wildcard has to expand in a root shell. The entry for the running kernel holds its title, kernel, initramfs and options; the grub_* lines are GRUB extensions. The last command is worth remembering: the running kernel is 211.16.1, but grubby reports 211.60.1 as the default, because an update installed a newer kernel that this VM has not started yet. A difference between these two lines usually means a kernel update is waiting for a reboot, but a default someone pinned with grubby --set-default or a one-off boot of another entry looks the same, so confirm it:
dnf needs-restarting -r compares what is running with what is installed and exits 1 when a reboot is needed, here because of kernel-core, so it also works in scripts and monitoring.
When a kernel package is installed, kernel-install runs the plugins in /usr/lib/kernel/install.d/: 20-grub.install writes the BLS entry and 50-dracut.install builds the initramfs, and /etc/kernel/cmdline is where kernel-install looks first for the options. To change the command line on RHEL, use grubby --update-kernel=ALL --args=..., as the method lesson did with psi=1; on Ubuntu, add a drop-in to /etc/default/grub.d/ and run update-grub, as the exercise at the end does. Either way the change takes effect at the next boot.
The kernel and the dracut initramfs
The kernel's own messages, with timestamps in seconds since it started, show the rest of the handover. -t kernel -t systemd selects the messages of the kernel and of systemd.
At 0.20 s the kernel prints the command line it received. It then unpacks the initramfs, a compressed archive GRUB loaded into memory next to the kernel, frees it once unpacked, and runs /init from it. /init is systemd, which reports "Running in initrd". The initramfs exists because what the kernel needs to find the root filesystem (storage drivers built as modules, LVM, software RAID, LUKS encryption, network storage) lives on disks it cannot read yet. At 1.31 s systemd switches the root to the real filesystem and starts again from there.
Both distributions build the initramfs with dracut: Ubuntu 26.04 uses dracut 110, and update-initramfs, the command older Ubuntu guides use, is now a wrapper shipped by the dracut package (initramfs-tools is not installed). lsinitrd lists the image's contents; /init is a link to systemd. RHEL 10 uses dracut 107 with the same layout and names the file initramfs-VERSION.img instead of initrd.img-VERSION.
Kernel packages rebuild the image on install. Rebuild it yourself after changing something it contains, such as a storage driver option, /etc/crypttab or a file in /etc/dracut.conf.d/: sudo update-initramfs -u on Ubuntu, sudo dracut -f on RHEL. A broken image stops the boot in the initramfs, so keep the previous kernel and its image installed until the new one has booted.
systemd takes over: timing the boot
systemd-analyze time splits the boot into the kernel (117 ms), the initrd (1.248 s) and userspace (2.884 s), and says when the default target was reached. The default target is graphical.target because that is what this cloud image sets (systemctl get-default); on a server with no display manager it adds nothing to multi-user.target, and the result is the same. There are no firmware or loader figures: the boot loader has to report them through the Boot Loader Interface, which systemd-boot implements and GRUB does not, so a GRUB machine never shows them. The same command on RHEL tells a different story.
multi-user.target was reached after 2.95 s, yet startup "finished" after 6.9 s of userspace. systemd reports startup as finished only when every job queued at boot has completed, and cloud-final.service, a one-shot cloud-init job that runs at the end of the boot, took 3.8 s and finished last. The machine was usable before that. dnf-makecache.service at the top of the list did not delay the boot at all: its timer started it half an hour later. Read both lines of systemd-analyze time before calling a boot slow.
systemd-analyze blame ranks units by how long they took to start, and that ranking alone does not tell you what delayed the boot. It lists every unit started since boot, including timer jobs that run hours later, and here its top entries are device units, none of which is on the path below. A slow unit that nothing waits for does not delay the boot. critical-chain shows the path that did gate the default target. After @ is the time, counted from the start of userspace, at which the unit became active or started; after + is how long it took to start. Units that ran in the initrd, such as systemd-fsck-root.service and dracut-pre-mount.service, have no times on this chain. sshd-vsock.socket is the SSH socket the VM tool adds (the units lesson explains it). Here snapd.seeded.service held up multi-user.target for 541 ms. Optimise the units on this chain; shortening anything else moves nothing.
The journal keeps each boot separately (Ubuntu stores it persistently in /var/log/journal). -b -1 is the previous boot, and its last lines show how it ended: journald received SIGTERM from systemd-shutdown and stopped, an orderly shutdown. A crash or power loss ends mid-stream with no shutdown messages. After any unexpected reboot, systemctl --failed lists the units that did not start.
When the boot stops
Without a login you work from the console, which on a cloud instance is the provider's serial console. At the GRUB menu, press e on an entry, add parameters to the end of the linux line and press Ctrl-X or F10 to boot once with them; nothing is saved. The useful ones: systemd.unit=rescue.target gives a single-user shell with the local filesystems mounted and only basic services; systemd.unit=emergency.target gives a shell with nothing else started, not even other mounts; rd.break (dracut) stops in the initramfs before the switch to the real root; and init=/bin/bash replaces systemd with a shell as the last resort, with no services, no logging and no clean shutdown.
Rescue and emergency mode ask for root's password through sulogin, and on Ubuntu 26.04, as on the Rocky cloud image used here, the root account is locked (the L). sulogin then prints "Cannot open access to console, the root account is locked." and no shell opens. Adding systemd.setenv=SYSTEMD_SULOGIN_FORCE=1 to the same boot makes it start the shell without a password, which is exactly why console access and the boot menu deserve protection (on RHEL, grub2-setpassword). The classic cause of a remote server stopping in emergency mode is a local /etc/fstab entry that cannot be mounted, such as a mistyped UUID: local-fs.target fails and systemd starts emergency.target. The essentials lesson "Disks, filesystems and mounts" shows the safe routine for editing /etc/fstab that prevents it.
/etc/fstab only when you can reach the console if the next boot fails, keep the previous kernel installed, and schedule the reboot while you are watching. Code that runs this early also runs before most monitoring, which is why the advanced security course (optional) treats boot-time persistence separately.Try this
On Ubuntu, add a harmless parameter with a drop-in and prove where it lands: echo 'GRUB_CMDLINE_LINUX_DEFAULT="$GRUB_CMDLINE_LINUX_DEFAULT loglevel=4"' | sudo tee /etc/default/grub.d/99-sd-boot-lab.cfg, then sudo update-grub. Expect update-grub to list each file it sources, including yours, and each kernel it finds. sudo grep -m1 -E '^\s+linux\s' /boot/grub/grub.cfg now ends in loglevel=4, while cat /proc/cmdline does not, because the running kernel booted before the change. Undo it with sudo rm /etc/default/grub.d/99-sd-boot-lab.cfg and sudo update-grub, and check that the linux line is back to what it was. Do not reboot a machine you cannot reach at its console.
Takeaway
Name the stage before you fix anything: the firmware entry (efibootmgr, bootctl), the boot loader and its command line (grub.cfg or BLS entries against /proc/cmdline), the initramfs (lsinitrd, journalctl -k), or systemd (critical-chain); each leaves evidence you can read on a running machine.