Recipe: a Python web app

A slim, non-root, hash-locked image for a WSGI app.

Intermediate13 min · lesson 18 of 24
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (6 files, 5 KB): r-python.tar.gz. The lab VM shares no folders with your computer, so fetch them inside the VM: cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-hard/r-python.tar.gz && tar -xzf r-python.tar.gz, which creates ~/lab/r-python/. SHA-256: c29eaf24232334b224154d4d04915dd70b1a99c5f945e788802aad51a4f4c5fd

This recipe packages a small Flask service with a hashed lock file, a virtual environment built in one stage and copied into a slim final stage, a numeric non-root user, and a health check that does not need curl. The case behind it is an image scan that flags gunicorn 22.0.0: CVE-2024-6827, an HTTP request-smuggling bug fixed in 23.0.0. The bump itself is one line. The follow-up work is finding which packages float unpinned with every build, what else the bump changed (a different Werkzeug, for example), and whether the build would notice a mirror serving a modified wheel.

Everything runs on the main lab VM, secopslog-docker, as ubuntu in ~/lab/r-python. The lesson files download has the app, the requirements files, the Dockerfile, .dockerignore and compose.yaml. The BuildKit mechanics used here (stages, COPY --from, cache mounts) are explained in "BuildKit builds: stages, cache and mounts"; this lesson applies them to Python.

The app and its lock

app.py
import os
import socket
from flask import Flask, jsonify
app = Flask(__name__)
@app.get("/health")
def health():
return jsonify(status="ok")
@app.get("/")
def index():
# print() writes to stdout; the request shows up in `docker logs` only if stdout is unbuffered
print(f"served / from {socket.gethostname()}")
return jsonify(service="orders", worker_pid=os.getpid())

/health is for the health check; / answers with the worker's process ID and prints a line, which the buffering section uses later. The dependencies come in two files. requirements.in names what the code imports, pinned to exact releases; requirements.txt is generated from it and pins everything those packages pull in, each with the SHA-256 hashes of the files PyPI publishes for that release:

ubuntu@secopslog-docker:~/lab/r-python · Docker 29.8.2
$ cat requirements.in; grep -c -- "--hash=sha256" requirements.txt; grep -A2 "^gunicorn==" requirements.txt
# Direct dependencies only. requirements.txt is generated from this file. flask==3.1.3 gunicorn==26.2.0 103 gunicorn==26.2.0 \ --hash=sha256:62b864895d9ebff0b2f9867ba04fe811c93121596540830c9c916d0769668447 \ --hash=sha256:bd249d0b3f7972f7432f0a6b6ff3b3ee2d129f70cd1ff6c09a9dd9e29a2b88e3

Two direct dependencies became eight pinned packages with 103 hashes between them. Most packages have two hashes (a wheel and a source archive); markupsafe has dozens because it ships a compiled wheel per Python version, OS and CPU, and the lock lists all of them so the same file installs on arm64 and amd64. When pip sees hashes in a requirements file it switches to hash-checking mode: every file it downloads must match a listed hash, and every requirement must be pinned and hashed. A lock is only useful if it can be regenerated and reviewed, so the lab rebuilds it with uv in a throwaway container and compares:

ubuntu@secopslog-docker:~/lab/r-python · Docker 29.8.2
$ docker run --rm --user "$(id -u):$(id -g)" -e HOME=/tmp -v "$PWD":/src:ro -w /tmp python:3.14-slim sh -c ' pip install --quiet --disable-pip-version-check uv==0.12.23 && cp /src/requirements.in . && python -m uv pip compile --quiet --generate-hashes --python-version 3.14 \ --exclude-newer 2026-10-01 requirements.in -o requirements.txt && diff -u /src/requirements.txt requirements.txt && echo "requirements.txt matches requirements.in"'
... requirements.txt matches requirements.in

No diff, so the committed lock is exactly what requirements.in resolves to. --exclude-newer makes the resolution reproducible: uv ignores anything published after that date, so rerunning the command next month gives the same file until you move the date on purpose. To bump gunicorn, change requirements.in and the date, regenerate, and review the diff, which shows every transitive change the bump brings. pip-compile --generate-hashes from pip-tools produces the same kind of file if you prefer it. Now the protection itself, with a requirement whose hash does not match what PyPI serves:

ubuntu@secopslog-docker:~/lab/r-python · Docker 29.8.2
$ printf 'gunicorn==26.2.0 --hash=sha256:%064d\n' 0 > tampered.txt docker run --rm -v "$PWD/tampered.txt:/tmp/r.txt:ro" python:3.14-slim \ pip install --quiet --disable-pip-version-check --root-user-action=ignore --require-hashes --no-deps -r /tmp/r.txt
... ERROR: THESE PACKAGES DO NOT MATCH THE HASHES FROM THE REQUIREMENTS FILE. If you have updated the package versions, please update the hashes. Otherwise, examine the package contents carefully; someone may have tampered with them. gunicorn==26.2.0 from https://files.pythonhosted.org/packages/fe/85/7522a52e5e2f42faf1a129113ab63e548c42e103e9af395b7bfe65e403e2/gunicorn-26.2.0-py3-none-any.whl (from -r /tmp/r.txt (line 1)): Expected sha256 0000000000000000000000000000000000000000000000000000000000000000 Got bd249d0b3f7972f7432f0a6b6ff3b3ee2d129f70cd1ff6c09a9dd9e29a2b88e3

pip downloaded gunicorn 26.2.0, computed its hash and refused to install it. A tampered mirror, a compromised cache or a re-uploaded file fails the same way, before any of its code runs. The hashes say nothing about whether the original release is safe; that is what scanning and update tooling are for ("Pinning, SBOMs, provenance and scanning" (Advanced container security)).

The Dockerfile

Dockerfile
# syntax=docker/dockerfile:1
FROM python:3.14-slim AS build
ENV PIP_DISABLE_PIP_VERSION_CHECK=1
RUN python -m venv /opt/venv
COPY requirements.txt /tmp/requirements.txt
RUN --mount=type=cache,target=/root/.cache/pip \
/opt/venv/bin/pip install --require-hashes --only-binary=:all: -r /tmp/requirements.txt
FROM python:3.14-slim
ENV PATH=/opt/venv/bin:$PATH \
PYTHONUNBUFFERED=1 \
PYTHONDONTWRITEBYTECODE=1
COPY --from=build /opt/venv /opt/venv
WORKDIR /app
COPY app.py .
USER 10001:10001
EXPOSE 8000
HEALTHCHECK --interval=10s --timeout=3s --start-period=10s --retries=3 \
CMD ["python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=2)"]
CMD ["gunicorn", "--bind", "0.0.0.0:8000", "--workers", "2", "--access-logfile", "-", "--no-control-socket", "app:app"]

The build stage creates a virtual environment in /opt/venv and installs the lock into it. --require-hashes makes hash checking mandatory even if someone later adds an unhashed line. --only-binary=:all: forbids source distributions, so pip installs wheels or fails; it never falls back to compiling C code, which would need a compiler the slim image does not have and would make the build depend on whatever headers happen to be installed. The cache mount keeps pip's download cache between builds without putting it in a layer.

The final stage starts again from python:3.14-slim and copies only /opt/venv and app.py. pip's cache, the build stage's temporary files and the requirements file stay behind. Copying a venv between stages works because a venv is not self-contained: its python is a symlink to the base image's interpreter, here /usr/local/bin/python3.14. Build and run stages must therefore use the same base image (ideally the same digest); a venv built on python:3.13-slim and copied into a 3.14 image points at an interpreter that does not exist.

PATH puts the venv first, so gunicorn and python resolve there without an activate script. PYTHONUNBUFFERED=1 is covered below; PYTHONDONTWRITEBYTECODE=1 stops Python from trying to write .pyc files into directories the app user cannot write anyway. USER 10001:10001 is numeric: no user of that ID exists in the image and none is needed, Kubernetes runAsNonRoot can verify a number, and nothing in /app or /opt/venv belongs to it, so the process cannot modify its own code. "Run as non-root" (Advanced container security) covers choosing these IDs and matching volume ownership.

The slim image has no curl or wget, so the health check uses Python itself: urllib.request.urlopen raises on a refused connection or an HTTP error status, which makes the command exit non-zero. Gunicorn binds 0.0.0.0 because a server bound to 127.0.0.1 inside the container is unreachable through a published port. --no-control-socket switches off a feature new in gunicorn 26: a Unix control socket the master creates under $XDG_RUNTIME_DIR or $HOME/.gunicorn. A numeric user with no passwd entry gets HOME=/, which the app user cannot write, so without the flag gunicorn logs Control server error: [Errno 13] Permission denied: '/.gunicorn' at startup (a first run of this lab did); a container is managed with signals anyway. Two workers suit the lab; for real traffic start near two per CPU core available to the container and measure. For an ASGI app (FastAPI, Starlette) swap the command for uvicorn with --workers, or keep gunicorn with uvicorn worker processes; the rest of the recipe stays the same.

The build context is kept small with an allow-list instead of a deny-list:

.dockerignore
# Only the Dockerfile's COPY sources are needed; keep everything else out of the build context
**
!app.py
!requirements.txt

Everything is excluded except the two files the Dockerfile copies, so a local .venv built for another OS, .git, __pycache__ or an .env with credentials cannot end up in the image by accident. When you add a file to a COPY, add it here too; the build fails loudly if you forget.

Build and run

compose.yaml
name: lab-orders
services:
web:
build: .
image: lab-orders-web:1
ports:
- "127.0.0.1:8000:8000"
restart: unless-stopped

Compose builds the image, names it lab-orders-web:1, and publishes port 8000 on the VM's loopback address only. --no-cache shows a first, uncached build:

ubuntu@secopslog-docker:~/lab/r-python · Docker 29.8.2
$ docker compose build --no-cache
... #9 [build 2/4] RUN python -m venv /opt/venv #9 DONE 3.4s #10 [build 3/4] COPY requirements.txt /tmp/requirements.txt #10 DONE 0.1s #11 [build 4/4] RUN --mount=type=cache,target=/root/.cache/pip /opt/venv/bin/pip install --require-hashes --only-binary=:all: -r /tmp/requirements.txt ... #11 15.56 Collecting blinker==1.9.0 (from -r /tmp/requirements.txt (line 3)) ... #11 35.95 Downloading blinker-1.9.0-py3-none-any.whl (8.5 kB) #11 36.10 Collecting click==8.5.0 (from -r /tmp/requirements.txt (line 7)) #11 36.14 Downloading click-8.5.0-py3-none-any.whl (125 kB) #11 36.46 Collecting flask==3.1.3 (from -r /tmp/requirements.txt (line 11)) #11 36.52 Downloading flask-3.1.3-py3-none-any.whl (103 kB) #11 36.77 Collecting gunicorn==26.2.0 (from -r /tmp/requirements.txt (line 15)) #11 36.83 Downloading gunicorn-26.2.0-py3-none-any.whl (228 kB) #11 37.07 Collecting itsdangerous==2.2.0 (from -r /tmp/requirements.txt (line 19)) #11 37.11 Downloading itsdangerous-2.2.0-py3-none-any.whl (16 kB) #11 37.24 Collecting jinja2==3.1.6 (from -r /tmp/requirements.txt (line 23)) #11 37.28 Downloading jinja2-3.1.6-py3-none-any.whl (134 kB) #11 37.63 Collecting markupsafe==3.0.3 (from -r /tmp/requirements.txt (line 27)) #11 37.68 Downloading markupsafe-3.0.3-cp314-cp314-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl (24 kB) #11 37.75 Collecting werkzeug==3.1.9 (from -r /tmp/requirements.txt (line 121)) #11 37.79 Downloading werkzeug-3.1.9-py3-none-any.whl (228 kB) #11 37.99 Installing collected packages: markupsafe, itsdangerous, gunicorn, click, blinker, werkzeug, jinja2, flask #11 38.77 #11 38.77 Successfully installed blinker-1.9.0 click-8.5.0 flask-3.1.3 gunicorn-26.2.0 itsdangerous-2.2.0 jinja2-3.1.6 markupsafe-3.0.3 werkzeug-3.1.9 #11 DONE 39.0s ... Image lab-orders-web:1 Built

Step #11 is the hash-checked install: eight wheels, including the aarch64 build of markupsafe (an amd64 machine downloads the x86_64 wheel from the same lock). The elided lines are BuildKit housekeeping and pip Retrying warnings, which also account for most of the 39 seconds. They come from the recording VM's 1280-byte link MTU (a VPN) meeting Docker's 1500-byte bridge; the lab README's troubleshooting section has the fix, and on a normal network the install takes a few seconds.

ubuntu@secopslog-docker:~/lab/r-python · Docker 29.8.2
$ docker compose up -d --wait
... Container lab-orders-web-1 Waiting Container lab-orders-web-1 Healthy
$ docker compose ps
NAME IMAGE COMMAND SERVICE CREATED STATUS PORTS lab-orders-web-1 lab-orders-web:1 "gunicorn --bind 0.0…" web 7 seconds ago Up 6 seconds (healthy) 127.0.0.1:8000->8000/tcp
$ curl -s http://127.0.0.1:8000/health; echo curl -s http://127.0.0.1:8000/; echo
{"status":"ok"} {"service":"orders","worker_pid":6}
$ docker image ls --format '{{.Repository}}:{{.Tag}} {{.Size}}' | grep -E '^(lab-orders-web:1|python:3.14-slim) '
lab-orders-web:1 233MB python:3.14-slim 206MB

--wait returned once the image's HEALTHCHECK reported healthy; Compose uses the image's health check when the service defines none. The worker PID varies between runs and requests. Sizes vary with the base image version: the app adds about 27 MB on top of python:3.14-slim, almost all of it the venv.

What is inside the container

ubuntu@secopslog-docker:~/lab/r-python · Docker 29.8.2
$ docker compose exec -T web id docker compose exec -T web sh -c 'for t in curl wget gcc; do command -v $t || echo "$t: not found"; done' docker compose exec -T web sh -c 'echo "HOME=$HOME"; readlink -f /opt/venv/bin/python; command -v gunicorn'
uid=10001 gid=10001 groups=10001 curl: not found wget: not found gcc: not found HOME=/ /usr/local/bin/python3.14 /opt/venv/bin/gunicorn
$ docker inspect -f '{{.State.Health.Status}}, last check exit {{(index .State.Health.Log 0).ExitCode}}, {{len .State.Health.Log}} checks logged' lab-orders-web-1
healthy, last check exit 0, 1 checks logged
$ docker compose logs --no-log-prefix web
[2026-10-07 20:03:39 +0000] [1] [INFO] Starting gunicorn 26.2.0 [2026-10-07 20:03:39 +0000] [1] [INFO] Listening at: http://0.0.0.0:8000 (1) [2026-10-07 20:03:39 +0000] [1] [INFO] Using worker: sync [2026-10-07 20:03:39 +0000] [6] [INFO] Booting worker with pid: 6 [2026-10-07 20:03:40 +0000] [7] [INFO] Booting worker with pid: 7 127.0.0.1 - - [07/Oct/2026:20:03:44 +0000] "GET /health HTTP/1.1" 200 16 "-" "Python-urllib/3.14" 172.18.0.1 - - [07/Oct/2026:20:03:45 +0000] "GET /health HTTP/1.1" 200 16 "-" "curl/8.18.0" served / from 99de8a603fdb 172.18.0.1 - - [07/Oct/2026:20:03:45 +0000] "GET / HTTP/1.1" 200 36 "-" "curl/8.18.0"

The process runs as 10001 with no name and no supplementary groups, and HOME is / because no passwd entry exists for that ID. There is no curl, wget or compiler. The venv's python resolves to the base image's interpreter, which is the dependency described above. The venv still contains its own pip; delete /opt/venv/bin/pip* at the end of the build stage if the running image should not be able to install packages. The health check passed and Docker keeps the last five results. The logs show gunicorn's own messages, the access log lines from --access-logfile -, and the served / lines from print().

Slim, Alpine or distroless

Older advice says Python on Alpine is slow to build because PyPI has only glibc wheels. That is out of date: musllinux wheels (PEP 656) exist for most popular packages, including NumPy, pandas and psycopg's binary package. The remaining arguments against Alpine for Python are narrower. Some packages still publish only manylinux wheels, and with --only-binary=:all: the build then fails instead of compiling. musl's DNS resolver and memory allocator behave differently from glibc's, which shows up under load. And the slim image is usually only a few tens of megabytes larger. Use slim unless you have measured a reason not to. A distroless Python base removes the shell and package manager as well; "Minimal bases: distroless, scratch, static and Alpine" (Advanced container security) covers the trade-offs and how to debug such images.

PYTHONUNBUFFERED

When stdout is a pipe, which it is under Docker without -t, Python buffers print() output in blocks of several kilobytes. Gunicorn's own messages go to stderr through the logging module, which flushes each record, so they appear in docker logs either way. To see the difference, run the image with the variable set to an empty string, which turns the setting off, and without the access log:

ubuntu@secopslog-docker:~/lab/r-python · Docker 29.8.2
$ docker run -d --name lab-orders-buffered -e PYTHONUNBUFFERED= -p 127.0.0.1:8001:8000 \ lab-orders-web:1 gunicorn --bind 0.0.0.0:8000 --no-control-socket app:app until [ "$(docker inspect -f '{{.State.Health.Status}}' lab-orders-buffered)" = healthy ]; do sleep 1; done curl -s http://127.0.0.1:8001/; echo docker logs lab-orders-buffered 2>&1 | grep 'served /' || echo "no print() output in the log yet"
793215e4910639a86af41430e5c0b97ff2568fc275a267b9f27fcd9e03964c48 {"service":"orders","worker_pid":7} no print() output in the log yet
$ docker stop lab-orders-buffered docker logs lab-orders-buffered 2>&1 | grep 'served /' docker rm lab-orders-buffered
lab-orders-buffered served / from 793215e49106 lab-orders-buffered

The request was answered, and its print() line was nowhere in the log. It appeared only when docker stop made the worker exit and flush its buffer. Had the container been killed with SIGKILL (an OOM kill, docker kill), the line would be gone. The access log matters for the comparison: with --access-logfile -, gunicorn writes each access line to the same stdout stream and flushes it, which pushes earlier print() output out too, so a buffering problem can hide until someone turns the access log off. Set PYTHONUNBUFFERED=1 in the image, and send application messages through the logging module rather than print().

Clean up. --rmi all removes the image built for the service:

ubuntu@secopslog-docker:~/lab/r-python · Docker 29.8.2
$ docker compose down --rmi all rm -f tampered.txt
Container lab-orders-web-1 Stopping Container lab-orders-web-1 Stopped Container lab-orders-web-1 Removing Container lab-orders-web-1 Removed Network lab-orders_default Removing Network lab-orders_default Removed Image lab-orders-web:1 Removing Image lab-orders-web:1 Removed
Quick check
01A build of this recipe fails with THESE PACKAGES DO NOT MATCH THE HASHES FROM THE REQUIREMENTS FILE for gunicorn 26.2.0, although nobody changed requirements.txt. What does that tell you?
Incorrect — pip installs the pinned version happily; hash checking compares bytes, not release dates.
Incorrect — A wheel for the wrong platform is skipped, not reported as a hash mismatch; the lock lists hashes for every platform's file.
Incorrect — A corrupt cached file would also fail the check, but the message is the same for any byte change, and deleting the cache hides the question of where the bytes came from.
Correct — That is the purpose of the hashes: a modified or substituted file fails before its code runs. Find out where the different file came from.
02To save build time, a teammate changes only the build stage to FROM python:3.13-slim AS build and keeps the final stage on python:3.14-slim. What breaks?
Incorrect — A venv links to the base image's interpreter; the lab showed /opt/venv/bin/python resolving to /usr/local/bin/python3.14.
Correct — The venv's interpreter is a symlink into the build stage's base image, so both stages must use the same base.
Incorrect — No user is created in either stage; a numeric USER needs no passwd entry.
Incorrect — The lock lists distinct hashes per file; installing in 3.13 picks other wheels, but the failure is the missing interpreter in the final stage.
03A service from this recipe runs without an access log and with PYTHONUNBUFFERED set to an empty string by a platform default. Which symptom do you expect in docker logs?
Correct — The lab showed the served / line missing after the request and appearing only when docker stop let the worker flush its buffer.
Incorrect — Gunicorn's messages go to stderr and logging handlers flush each record; they appear regardless.
Incorrect — Buffering delays output; it does not duplicate it.
Incorrect — The variable affects Python's stdout and stderr, not sockets; the buffered container in the lab reported healthy.

Try this

Work through “PYTHONUNBUFFERED” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from recipe: a python web app, keep “PYTHONUNBUFFERED”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related