Sandboxes and runtime detection
gVisor, Kata and Falco: what each adds, what it costs, and when to use it.
Every container on a host shares one kernel. A default container's system calls run against the host kernel directly, so a kernel bug reachable from inside the container ("Capability and kernel escapes" has the examples) lands on the host. Two runtimes change that by putting a second kernel between the container and the host, and one tool watches for the behaviours an escape produces when prevention fails. This lesson installs gVisor and Falco on the sec VM, measures what each adds and costs, and shows where Kata Containers can and cannot run.
secopslog-docker-sec). The lab installs gVisor and Falco, registers a Docker runtime and loads a kernel probe. If the VM does not exist, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec on Lima), and reset it at any point with ./setup/create-lab.sh --profile sec --recreate. On the sec VM docker runs through sudo. Pull the two images the lesson uses first:gVisor: a user-space kernel
gVisor runs each container against the Sentry, a process that re-implements the Linux system-call interface in user space. When the container calls open or socket, the Sentry services it; only a small, fixed set of calls reaches the host kernel, and the Sentry itself runs under a tight seccomp filter. Install it from the official release and verify the checksum before trusting the binary. The downloads go in a working directory of their own (mkdir -p ~/lab/sandboxes && cd ~/lab/sandboxes), which the clean-up removes:
sha512sum -c checks the download against the published digest; runsc --version is release 20260928.0. runsc install registers a Docker runtime named runsc by merging it into /etc/docker/daemon.json. Back the file up first, as "Configuring the daemon safely" (Docker in depth) teaches, so the clean-up can put it back exactly; then validate and reload:
Two lines of that output deserve a look. runsc install also downloaded a second tarball of "sidecar binaries" into /usr/local/bin/gvisor-bin, with no checksum shown, and says the feature is best-effort and will stop working (gVisor issue 13718 tracks it). The runtime itself does not need it here, and the clean-up removes it. configuration OK is dockerd --validate confirming the merged file before the reload. The daemon now lists the new runtime:
The daemon now lists runsc alongside runc. Run the same image under each and read the kernel it reports:
Under runc the container sees the host kernel; under runsc it sees the Sentry's own version. The dmesg banner says the same thing out loud:
dmesg and uname are reported by the Sentry, so they are easy for a workload to read but must never be your check for whether a container is sandboxed. A compromised container could print whatever it likes. Confirm the runtime from the host with docker inspect -f "{{.HostConfig.Runtime}}", which reads the daemon's record, not the container's.What the sandbox costs
The Sentry services syscalls in user space, so a syscall-heavy workload pays for every call. Time the same fixed workload under each runtime:
The numbers vary between runs and machines, but the ratio is the point: a workload that does little but make system calls runs several times slower under gVisor. The other cost is compatibility. gVisor implements most of the Linux syscall surface but not all of it, and returns ENOSYS for calls it has not built, so a program that reaches for an unimplemented or partly implemented kernel interface can fail to start (gVisor's compatibility notes list them). That is why a sandbox is a targeted tool, not a default: use it for code you did not write and cannot vouch for, such as customer-supplied builds or CI jobs running arbitrary pull requests. For your own services, the controls from earlier in this course are the floor: a minimal image, non-root, dropped capabilities, seccomp and a read-only root filesystem. A sandbox is what you add when a shared kernel is a line you cannot afford to cross.
Falco: watch the behaviour when prevention fails
A sandbox shrinks the attack surface and detects nothing. Someone will find a bug in the Sentry, or you will keep a workload on plain runc because the sandbox broke it. So you also watch for the behaviours an escape produces and alert when one appears. Falco loads a probe into the host kernel (the modern eBPF driver, bundled in the binary) and checks the syscall stream against rules. That probe sees containers whose syscalls reach the host kernel, which means runc containers. Under runsc the Sentry services the workload's syscalls in user space, so the probe sees only the Sentry's own small set of host calls, not the workload's execve or connect; Falco needs its separate gVisor event source, fed by runsc's trace points, to see inside. Under Kata the syscalls go to the guest kernel, so Falco has to run inside the guest. Plan detection per runtime. Download the build for your CPU (uname -m prints x86_64 or aarch64, the names Falco's downloads use), and check both the key and the signature:
The script compares the key's fingerprint with Falco's published packaging key, 478B2FBBC75F4237B731DA4365106822B35B1B1F, and gpg --verify confirms the tarball was signed with that key. [unknown] is gpg saying you have not marked the key as trusted, which the fingerprint check covers. Install and run it with the modern eBPF driver:
Falco runs in the background and writes its log and its alerts to /tmp/falco.out. The loop waits until Falco reports that it opened the syscall source with the modern BPF probe, which is the moment it starts seeing events. The reason it lands so well on containers you have already hardened is that the behaviours worth alerting on are behaviors that should be impossible in a tight container. Open a shell in a running container, the classic first move after a foothold, and the stock "Terminal shell in container" rule fires:
The alert fired the moment docker exec started a shell with a terminal attached: process=sh, parent=containerd-shim, the command, and terminal=34816 (a non-zero TTY is what this rule keys on). The full line goes on to name the container and image (container_name=lab-payments, nginx:1.30-alpine); the sed trims it for width. A distroless, non-root, read-only container ships no shell at all, so a shell starting in one is an alarm rather than noise. Every control you added earlier in this course doubles as a tripwire here. Stop Falco and remove the container:
file_output in falco.yaml. And do not over-trust any single rule. The stock "Terminal shell in container" rule requires an attached terminal (proc.tty != 0), so a reverse shell that pipes a raw socket into /bin/sh without a pseudo-terminal does not trip it. The stable ruleset does cover that case with other rules, "Redirect STDOUT/STDIN to Network Connection in Container" and "Netcat Remote Code Execution in Container", so Falco is not blind to reverse shells; the gap is specific to the TTY-based shell rule, and the fix is to keep the redirect rule enabled rather than leaning on the terminal rule alone.Kata Containers: a guest kernel in a micro-VM
Kata goes further than gVisor: it wraps each container in a lightweight virtual machine with its own guest kernel, so an attacker must break out of the container and then out of a VM before the host is in reach. The cost is that a VM needs hardware virtualisation. Check for it on this lab VM:
There is no /dev/kvm here. The SecOpsLog lab VMs are themselves virtual machines and usually have no nested virtualisation (Apple Silicon Macs offer none), so Kata cannot run in this lab, and nothing below was executed here. On a bare-metal host (or a cloud instance with nested virtualisation) that has /dev/kvm, Kata installs as a containerd shim, io.containerd.kata.v2. There is no bare --runtime=kata unless you register that name yourself. Install Kata (its kata-manager script unpacks the release under /opt/kata and symlinks containerd-shim-kata-v2), then either call the shim by its fully qualified name:
docker run --runtime io.containerd.kata.v2 --rm alpine uname -r
or register a short name in daemon.json, the same mechanism runsc install used for gVisor, pointing at that shim type:
{"runtimes": {"kata": {"runtimeType": "io.containerd.kata.v2"}}}
Clean up: restore the daemon.json backup, which removes the runsc runtime entry, and restart the daemon (a reload left the old runtime listed in this lab), then remove the gVisor and Falco files, including the sidecar directory and Falco's configuration and rules:
The runtime list is back to the default. Falco's probe went with the process. For a fully clean VM, recreate it with ./setup/create-lab.sh --profile sec --recreate.
--runtime=runsc. Where does their code run?Try this
Work through “Kata Containers: a guest kernel in a micro-VM” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from sandboxes and runtime detection, keep “Kata Containers: a guest kernel in a micro-VM”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.