Containers vs virtual machines

One shared kernel versus one kernel per guest, and what that costs.

Beginner9 min · lesson 2 of 14

Run uname -sr inside an Alpine Linux container on the lab VM and it prints Linux 7.0.0-34-generic. That is not an Alpine kernel. It is the Ubuntu VM's own kernel, the same one uname reports outside the container. That one line explains most of the difference between containers and virtual machines, including why containers start so fast and why their isolation needs care.

Two ways to isolate a workload

A virtual machine (VM) is built by a hypervisor such as KVM, Hyper-V or Apple's Virtualization framework. The hypervisor presents virtual CPUs, memory, disks and network cards, and a complete guest operating system boots on them, with its own kernel. The kernel is the core of the operating system: it schedules processes, manages memory and talks to hardware. Every VM pays for its own kernel and its own init system before your application starts.

A container boots nothing. It is an ordinary process on the host, started by the container runtime with two kinds of kernel features around it. Namespaces limit what the process can see (its own process list, network interfaces, mount table, hostname). Control groups (cgroups) limit what it can use (CPU, memory, number of processes). The files it sees come from the image. The kernel underneath is the host's.

The two stacks
Virtual machines
App A + libraries
Guest OS + guest kernel
one per VM
Hypervisor
virtual hardware
Host hardware
Containers
App A + libraries
from the image
App B + libraries
from another image
Container runtime
namespaces + cgroups
One host kernel
shared by every container
A VM carries a whole operating system under each application. Containers share the kernel that is already running.

One kernel, different userlands

Compare the VM and a container side by side:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ uname -sr
Linux 7.0.0-34-generic
$ docker run --rm alpine:3.22 uname -sr
Linux 7.0.0-34-generic
$ grep PRETTY_NAME /etc/os-release
PRETTY_NAME="Ubuntu 26.04.1 LTS"
$ docker run --rm alpine:3.22 grep PRETTY_NAME /etc/os-release
PRETTY_NAME="Alpine Linux v3.22"

The kernel lines are identical. The /etc/os-release lines are not, because that file comes from the image. A "Debian" or "Alpine" container is that distribution's userland (its shell, libraries and package manager) running on whatever kernel the host has. Two consequences follow. A Linux container cannot run Windows programs, since there is no Windows kernel to call. And software that needs a particular kernel feature depends on the host, not on the image: an image cannot bring a newer kernel with it. Windows containers do exist, but they are a separate world that runs on Windows hosts with the Windows kernel.

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ systemd-detect-virt
qemu

systemd-detect-virt answers qemu: the lab "host" is itself a virtual machine. That stack is normal. Cloud instances are VMs, and Docker Desktop on a Mac runs a Linux VM, so most production containers run inside a VM. The containers share the VM's kernel, not the laptop's or the hypervisor's.

A container is a process you can see from the host

Start a container that does nothing for ten minutes, then look at it from both sides. The docker inspect -f template prints one field of the container's state, here its process ID on the host ("Inspect, logs and exec" covers templates). docker exec runs an extra command inside a running container, which gives the view from inside:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-idle alpine:3.22 sleep 600
6c0a2c47576e39834dc1aa3cb15c4acf32dbff22591cbd9bf39d3bc4493bc3cb
$ ps -o pid,user,args -p "$(docker inspect -f '{{.State.Pid}}' lab-idle)"
PID USER COMMAND 117155 root sleep 600
$ docker exec lab-idle ps
PID USER TIME COMMAND 1 root 0:00 sleep 600 7 root 0:00 ps

From the VM, the container is process 117155, a plain sleep 600 owned by root. Inside, the same process is PID 1 and it can see only itself and the ps you just ran. That is the PID namespace: the container gets its own numbering and its own, much shorter, process list. The PID on your VM will differ.

Each process has a set of namespace links under /proc/self/ns. Compare two of them:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ readlink /proc/self/ns/pid /proc/self/ns/user
pid:[4026531836] user:[4026531837]
$ docker exec lab-idle sh -c 'readlink /proc/self/ns/pid; readlink /proc/self/ns/user'
pid:[4026532321] user:[4026531837]

The PID namespace numbers differ, so the container really has its own. The user namespace number is the same on both sides. Docker does not create a user namespace by default, which means UID 0 in the container is UID 0 on the host kernel. The isolation you get by default covers process IDs, network, mounts, hostname and IPC, plus a cgroup namespace; user IDs are shared unless you opt into rootless Docker or user-namespace remapping. Advanced container security spends several lessons on what root in a container can and cannot do; for now, do not read "it runs in a container" as "it is not root".

What a container costs

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ time docker run --rm alpine:3.22 true
real 0m0.284s user 0m0.016s sys 0m0.016s
$ docker stats --no-stream lab-idle
CONTAINER ID NAME CPU % MEM USAGE / LIMIT MEM % NET I/O BLOCK I/O PIDS 6c0a2c47576e lab-idle 0.00% 584KiB / 5.76GiB 0.01% 486B / 84B 0B / 0B 1
$ docker image ls alpine:3.22
IMAGE ID DISK USAGE CONTENT SIZE EXTRA alpine:3.22 5291449c3df7 13.4MB 4.21MB U

With the image already on disk, creating a container, running true in it and deleting it took well under half a second of wall-clock time (real, 0.28 s in this run). There is no boot: the runtime sets up namespaces and cgroups and executes one program. The idle container uses under 600 KiB of memory and one process (PIDS 1), and docker stats shows the limit as the VM's whole memory because no limit was set. The image is 4.21 MB of compressed content and 13.4 MB on disk once unpacked; "Inspecting, exporting and cleaning up images" in Docker in depth explains those two columns. A VM, by comparison, holds memory for its guest kernel and services before your application starts, and it boots before it can do anything. Your timings and memory figures will differ a little.

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker rm -f lab-idle
lab-idle

The trade-off

Sharing a kernel is the price of that speed. A hypervisor puts a hardware-enforced boundary between guests; a container boundary is made of kernel features, and every container makes system calls into the same kernel. A kernel vulnerability can therefore cross from one container to the host and to its neighbours. That is why container hardening exists (dropping capabilities, seccomp filters, running as a non-root user, keeping the host kernel patched) and why some platforms put each container in a small VM or a user-space kernel, as with Kata Containers or gVisor.

Choose a VM when the workload needs a different kernel or operating system, when you must separate tenants that do not trust each other, or when the software expects to own the whole machine (a kernel module, a network appliance). Choose containers for many Linux services that can share a kernel, start in under a second and are rebuilt from an image on every release. Most teams run both, with containers inside VMs, and that combination is where this track's lab lives too.

Quick check
01An Alpine container on an Ubuntu 26.04 host runs uname -r. What does it print?
Incorrect — Images do not contain a running kernel. Even if a kernel package were installed in the image, the container would not boot it.
Incorrect — Docker assigns names and IDs, not kernels. There is no per-container kernel to report.
Correct — The container is a process on the host kernel; only /etc/os-release and the other userland files come from the Alpine image.
Incorrect — uname is a normal program in Alpine's BusyBox. It asks the kernel, and the kernel answers with its own version.
02On the lab VM, readlink /proc/self/ns/user gives the same number on the host and inside a container started with plain docker run. What does that tell you?
Correct — Docker leaves user namespaces off by default; only rootless mode or userns-remap maps container root to an unprivileged host user.
Incorrect — The PID namespace is a separate link, and it differed. Each namespace type is created independently.
Incorrect — /proc/self in the container is its own process's view. The pid link it printed was different, so it was not reading the host's.
Incorrect — The container ran as root (ps showed root). Sharing the user namespace means the UIDs are the same numbers, not the same account as your shell.
03A team wants to run customer-supplied plugins next to each other on one host, and the customers do not trust each other. What does this lesson suggest?
Incorrect — Namespaces are kernel features, not hardware boundaries. All the containers still share one kernel and its bugs.
Incorrect — Kernel vulnerabilities are exactly the risk of a shared kernel; namespaces do not protect the kernel from the process calling it.
Incorrect — Images do not ship a running kernel. Every container on the host uses the host's kernel.
Correct — Mutually untrusted tenants are the case where the hardware boundary of a VM, or a runtime such as Kata or gVisor, earns its cost.

Try this

Work through “The trade-off” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from containers vs virtual machines, keep “The trade-off”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related