Recipe: a Python web app
A slim, non-root, hash-locked image for a WSGI app.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-hard/r-python.tar.gz && tar -xzf r-python.tar.gz, which creates ~/lab/r-python/. SHA-256: c29eaf24232334b224154d4d04915dd70b1a99c5f945e788802aad51a4f4c5fdThis recipe packages a small Flask service with a hashed lock file, a virtual environment built in one stage and copied into a slim final stage, a numeric non-root user, and a health check that does not need curl. The case behind it is an image scan that flags gunicorn 22.0.0: CVE-2024-6827, an HTTP request-smuggling bug fixed in 23.0.0. The bump itself is one line. The follow-up work is finding which packages float unpinned with every build, what else the bump changed (a different Werkzeug, for example), and whether the build would notice a mirror serving a modified wheel.
Everything runs on the main lab VM, secopslog-docker, as ubuntu in ~/lab/r-python. The lesson files download has the app, the requirements files, the Dockerfile, .dockerignore and compose.yaml. The BuildKit mechanics used here (stages, COPY --from, cache mounts) are explained in "BuildKit builds: stages, cache and mounts"; this lesson applies them to Python.
The app and its lock
import osimport socketfrom flask import Flask, jsonifyapp = Flask(__name__)@app.get("/health")def health():return jsonify(status="ok")@app.get("/")def index():# print() writes to stdout; the request shows up in `docker logs` only if stdout is unbufferedprint(f"served / from {socket.gethostname()}")return jsonify(service="orders", worker_pid=os.getpid())
/health is for the health check; / answers with the worker's process ID and prints a line, which the buffering section uses later. The dependencies come in two files. requirements.in names what the code imports, pinned to exact releases; requirements.txt is generated from it and pins everything those packages pull in, each with the SHA-256 hashes of the files PyPI publishes for that release:
Two direct dependencies became eight pinned packages with 103 hashes between them. Most packages have two hashes (a wheel and a source archive); markupsafe has dozens because it ships a compiled wheel per Python version, OS and CPU, and the lock lists all of them so the same file installs on arm64 and amd64. When pip sees hashes in a requirements file it switches to hash-checking mode: every file it downloads must match a listed hash, and every requirement must be pinned and hashed. A lock is only useful if it can be regenerated and reviewed, so the lab rebuilds it with uv in a throwaway container and compares:
No diff, so the committed lock is exactly what requirements.in resolves to. --exclude-newer makes the resolution reproducible: uv ignores anything published after that date, so rerunning the command next month gives the same file until you move the date on purpose. To bump gunicorn, change requirements.in and the date, regenerate, and review the diff, which shows every transitive change the bump brings. pip-compile --generate-hashes from pip-tools produces the same kind of file if you prefer it. Now the protection itself, with a requirement whose hash does not match what PyPI serves:
pip downloaded gunicorn 26.2.0, computed its hash and refused to install it. A tampered mirror, a compromised cache or a re-uploaded file fails the same way, before any of its code runs. The hashes say nothing about whether the original release is safe; that is what scanning and update tooling are for ("Pinning, SBOMs, provenance and scanning" (Advanced container security)).
The Dockerfile
# syntax=docker/dockerfile:1FROM python:3.14-slim AS buildENV PIP_DISABLE_PIP_VERSION_CHECK=1RUN python -m venv /opt/venvCOPY requirements.txt /tmp/requirements.txtRUN --mount=type=cache,target=/root/.cache/pip \/opt/venv/bin/pip install --require-hashes --only-binary=:all: -r /tmp/requirements.txtFROM python:3.14-slimENV PATH=/opt/venv/bin:$PATH \PYTHONUNBUFFERED=1 \PYTHONDONTWRITEBYTECODE=1COPY --from=build /opt/venv /opt/venvWORKDIR /appCOPY app.py .USER 10001:10001EXPOSE 8000HEALTHCHECK --interval=10s --timeout=3s --start-period=10s --retries=3 \CMD ["python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=2)"]CMD ["gunicorn", "--bind", "0.0.0.0:8000", "--workers", "2", "--access-logfile", "-", "--no-control-socket", "app:app"]
The build stage creates a virtual environment in /opt/venv and installs the lock into it. --require-hashes makes hash checking mandatory even if someone later adds an unhashed line. --only-binary=:all: forbids source distributions, so pip installs wheels or fails; it never falls back to compiling C code, which would need a compiler the slim image does not have and would make the build depend on whatever headers happen to be installed. The cache mount keeps pip's download cache between builds without putting it in a layer.
The final stage starts again from python:3.14-slim and copies only /opt/venv and app.py. pip's cache, the build stage's temporary files and the requirements file stay behind. Copying a venv between stages works because a venv is not self-contained: its python is a symlink to the base image's interpreter, here /usr/local/bin/python3.14. Build and run stages must therefore use the same base image (ideally the same digest); a venv built on python:3.13-slim and copied into a 3.14 image points at an interpreter that does not exist.
PATH puts the venv first, so gunicorn and python resolve there without an activate script. PYTHONUNBUFFERED=1 is covered below; PYTHONDONTWRITEBYTECODE=1 stops Python from trying to write .pyc files into directories the app user cannot write anyway. USER 10001:10001 is numeric: no user of that ID exists in the image and none is needed, Kubernetes runAsNonRoot can verify a number, and nothing in /app or /opt/venv belongs to it, so the process cannot modify its own code. "Run as non-root" (Advanced container security) covers choosing these IDs and matching volume ownership.
The slim image has no curl or wget, so the health check uses Python itself: urllib.request.urlopen raises on a refused connection or an HTTP error status, which makes the command exit non-zero. Gunicorn binds 0.0.0.0 because a server bound to 127.0.0.1 inside the container is unreachable through a published port. --no-control-socket switches off a feature new in gunicorn 26: a Unix control socket the master creates under $XDG_RUNTIME_DIR or $HOME/.gunicorn. A numeric user with no passwd entry gets HOME=/, which the app user cannot write, so without the flag gunicorn logs Control server error: [Errno 13] Permission denied: '/.gunicorn' at startup (a first run of this lab did); a container is managed with signals anyway. Two workers suit the lab; for real traffic start near two per CPU core available to the container and measure. For an ASGI app (FastAPI, Starlette) swap the command for uvicorn with --workers, or keep gunicorn with uvicorn worker processes; the rest of the recipe stays the same.
The build context is kept small with an allow-list instead of a deny-list:
# Only the Dockerfile's COPY sources are needed; keep everything else out of the build context**!app.py!requirements.txt
Everything is excluded except the two files the Dockerfile copies, so a local .venv built for another OS, .git, __pycache__ or an .env with credentials cannot end up in the image by accident. When you add a file to a COPY, add it here too; the build fails loudly if you forget.
Build and run
name: lab-ordersservices:web:build: .image: lab-orders-web:1ports:- "127.0.0.1:8000:8000"restart: unless-stopped
Compose builds the image, names it lab-orders-web:1, and publishes port 8000 on the VM's loopback address only. --no-cache shows a first, uncached build:
Step #11 is the hash-checked install: eight wheels, including the aarch64 build of markupsafe (an amd64 machine downloads the x86_64 wheel from the same lock). The elided lines are BuildKit housekeeping and pip Retrying warnings, which also account for most of the 39 seconds. They come from the recording VM's 1280-byte link MTU (a VPN) meeting Docker's 1500-byte bridge; the lab README's troubleshooting section has the fix, and on a normal network the install takes a few seconds.
--wait returned once the image's HEALTHCHECK reported healthy; Compose uses the image's health check when the service defines none. The worker PID varies between runs and requests. Sizes vary with the base image version: the app adds about 27 MB on top of python:3.14-slim, almost all of it the venv.
What is inside the container
The process runs as 10001 with no name and no supplementary groups, and HOME is / because no passwd entry exists for that ID. There is no curl, wget or compiler. The venv's python resolves to the base image's interpreter, which is the dependency described above. The venv still contains its own pip; delete /opt/venv/bin/pip* at the end of the build stage if the running image should not be able to install packages. The health check passed and Docker keeps the last five results. The logs show gunicorn's own messages, the access log lines from --access-logfile -, and the served / lines from print().
Slim, Alpine or distroless
Older advice says Python on Alpine is slow to build because PyPI has only glibc wheels. That is out of date: musllinux wheels (PEP 656) exist for most popular packages, including NumPy, pandas and psycopg's binary package. The remaining arguments against Alpine for Python are narrower. Some packages still publish only manylinux wheels, and with --only-binary=:all: the build then fails instead of compiling. musl's DNS resolver and memory allocator behave differently from glibc's, which shows up under load. And the slim image is usually only a few tens of megabytes larger. Use slim unless you have measured a reason not to. A distroless Python base removes the shell and package manager as well; "Minimal bases: distroless, scratch, static and Alpine" (Advanced container security) covers the trade-offs and how to debug such images.
PYTHONUNBUFFERED
When stdout is a pipe, which it is under Docker without -t, Python buffers print() output in blocks of several kilobytes. Gunicorn's own messages go to stderr through the logging module, which flushes each record, so they appear in docker logs either way. To see the difference, run the image with the variable set to an empty string, which turns the setting off, and without the access log:
The request was answered, and its print() line was nowhere in the log. It appeared only when docker stop made the worker exit and flush its buffer. Had the container been killed with SIGKILL (an OOM kill, docker kill), the line would be gone. The access log matters for the comparison: with --access-logfile -, gunicorn writes each access line to the same stdout stream and flushes it, which pushes earlier print() output out too, so a buffering problem can hide until someone turns the access log off. Set PYTHONUNBUFFERED=1 in the image, and send application messages through the logging module rather than print().
Clean up. --rmi all removes the image built for the service:
THESE PACKAGES DO NOT MATCH THE HASHES FROM THE REQUIREMENTS FILE for gunicorn 26.2.0, although nobody changed requirements.txt. What does that tell you?FROM python:3.13-slim AS build and keeps the final stage on python:3.14-slim. What breaks?/opt/venv/bin/python resolving to /usr/local/bin/python3.14.PYTHONUNBUFFERED set to an empty string by a platform default. Which symptom do you expect in docker logs?served / line missing after the request and appearing only when docker stop let the worker flush its buffer.Try this
Work through “PYTHONUNBUFFERED” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from recipe: a python web app, keep “PYTHONUNBUFFERED”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.