Concurrency on Python 3.14: threads, processes, interpreters and asyncio
The GIL build and the free-threaded build, forkserver, and asyncio task groups.
tar -xzf scr-py-concurrency.tar.gz, which creates scr-py-concurrency/. SHA-256: 9c00d27e717ce040abdb828ea395d6e8beed0fdf25b6ea80a1d7761cc57a3ca8This lesson gives you measured answers to "should this run in threads, processes, interpreters or asyncio?" on Python 3.14. You will check which build of CPython you are running, time the same CPU-bound job four ways on the default build and on the free-threaded build, see C code that releases the GIL, learn how asyncio runs coroutines, fan out 200 slow HTTP calls with threads and with asyncio, stall an event loop on purpose, and watch a shared counter lose updates. Every number below was measured on the lab VM (4 vCPUs, Ubuntu 26.04). Timings change from run to run and from machine to machine; read them as ratios, and measure your own workload before you decide.
Refresher: "Functions, modules and exceptions" in py-sec owns if __name__ == "__main__":, and "Failure design" (the previous lesson) owns ExceptionGroup and except*. This lesson owns asyncio; the fleet lesson and the capstone build on it. Unpack the lesson files (the box at the top of this page) in your home directory and work in ~/scr-py-concurrency: they hold every script below, including slowserver.py, the small HTTP server the fan-out section starts and stops.
Two builds of CPython 3.14
The global interpreter lock (GIL) is a lock inside CPython that lets only one thread at a time run Python bytecode in one interpreter. It is a property of a build, not of the language. Python 3.14 comes in two builds: the default build, which keeps the GIL, and the free-threaded build, compiled with --disable-gil, where threads run Python code in parallel. PEP 703 introduced the free-threaded build in 3.13 as experimental; PEP 779 made it officially supported in 3.14, but it stays an optional, separate build (its binary is python3.14t), and nothing has been decided about making it the default. Ask the interpreter which one you have:
import sysimport sysconfigprint(sys.version)# Py_GIL_DISABLED is fixed when the interpreter is compiled; _is_gil_enabled() is the state now.print("free-threaded build:", sysconfig.get_config_var("Py_GIL_DISABLED") == 1)print("GIL enabled now: ", sys._is_gil_enabled())
Ubuntu's python3 is the default build: Py_GIL_DISABLED is not set and the GIL is on. /opt/cpython-3.14.7t is the lab's free-threaded build of the upstream 3.14.7 source; its version string says "free-threading build" and the GIL is off. Py_GIL_DISABLED is fixed when the interpreter is compiled; sys._is_gil_enabled() is the state right now, and PYTHON_GIL=1 (or -X gil=1) turns the GIL back on in a free-threaded build. A C extension that is not marked as safe without the GIL also turns it back on when imported, with a warning, so check at run time, not only at install time. The default build cannot go the other way: PYTHON_GIL=0 is a fatal error there.
Ubuntu 26.04 packages neither the free-threaded build nor 3.14.7. The archive has python3.14 at 3.14.4, and no package whose name looks like a free-threaded Python:
So the lab VM built upstream 3.14.7 from the python.org source twice when it was prepared: once with --disable-gil into /opt/cpython-3.14.7t, once without into /opt/cpython-3.14.7. These are the steps for the free-threaded one. A build takes several minutes on 4 CPUs, so the lab does not repeat it; the next terminal proves what was built:
# Build dependencies from the Ubuntu archivesudo apt-get install -y build-essential pkg-config libssl-dev zlib1g-dev libbz2-dev \libreadline-dev libsqlite3-dev libffi-dev liblzma-dev uuid-dev libgdbm-dev \libgdbm-compat-dev libzstd-dev# The source, checked against the SHA-256 that python.org publishes for this releasecurl -fsSLO https://www.python.org/ftp/python/3.14.7/Python-3.14.7.tar.xzecho '3b48dac8fb59f62eaa67ac83c1eb12bda1b7a08406dd286e252c11a66be27f81 Python-3.14.7.tar.xz' | sha256sum -ctar -xf Python-3.14.7.tar.xz && cd Python-3.14.7# Its own prefix: /usr/bin/python3 and everything Ubuntu installed with it stay untouched./configure --prefix=/opt/cpython-3.14.7t --disable-gilmake -j"$(nproc)"sudo make installsudo ln -sf python3.14t /opt/cpython-3.14.7t/bin/python3 # the same python3 name inside the prefix# The default build for comparison: a fresh copy of the source, ./configure --prefix=/opt/cpython-3.14.7
CONFIG_ARGS records the options an interpreter was configured with. Check all three:
The two /opt builds differ only in --disable-gil, which makes the timing comparison below fair. Ubuntu's own build does not use it either. Other routes to a free-threaded build: the python.org installers for macOS and Windows offer it as an optional component, and uv python install 3.14t downloads a prebuilt one (uv arrives in "Shipping automation" later in this course). If you have no free-threaded build, read the terminals that call /opt/cpython-3.14.7t; every other command runs on Ubuntu's python3.
CPU-bound Python: threads, processes and interpreters
The job is pure Python: generate host names and flag the random-looking ones by their character entropy, the way a DGA detector scores domains. score_chunk(seed) does 150,000 names and shares nothing with other calls:
"""CPU-bound work in pure Python: flag generated host names that look random (DGA-like)."""import mathimport randomLETTERS = "abcdefghijklmnopqrstuvwxyz0123456789"def entropy(name: str) -> float:counts: dict[str, int] = {}for ch in name:counts[ch] = counts.get(ch, 0) + 1n = len(name)return -sum(c / n * math.log2(c / n) for c in counts.values())def score_chunk(seed: int, count: int = 150_000) -> int:"""Generate count names from seed; return how many score above the threshold."""rng = random.Random(seed) # one generator per task: nothing is shared between workersflagged = 0for _ in range(count):name = "".join(rng.choices(LETTERS, k=16))if entropy(name) > 3.6:flagged += 1return flagged
compare.py runs 8 of those tasks serially, then on a pool of 4 threads, 4 processes and 4 interpreters. InterpreterPoolExecutor is new in 3.14: each worker is a separate interpreter inside the same process, with its own GIL. Before running it, see what the 3.14 default for processes on Linux does to a script without a __main__ guard:
from concurrent.futures import ProcessPoolExecutorfrom score import score_chunk# No __main__ guard: the fork server imports this file, so it would start a pool too.with ProcessPoolExecutor(max_workers=2) as pool:print(sum(pool.map(score_chunk, range(4))))
Since 3.14 the default start method on Linux is forkserver, not fork. With fork, each worker was a copy of the running program. With forkserver, the pool first starts a server process, which imports your main module once (spawn.import_main_path in the first traceback) so that the workers it forks from itself start with your functions loaded. Without the guard, your top-level code runs again inside that server and tries to start a pool of its own; multiprocessing stops it with the "bootstrapping phase" RuntimeError, and the server dies. The second traceback is the parent: it was talking to the server when the server exited, hence ConnectionResetError. (With spawn, the default on macOS and Windows, every worker imports the main module, so the guard is needed there too.) The fix is the guard, and the function a worker runs must be importable from a module, which is why score_chunk lives in score.py:
"""Run the same 8 CPU-bound tasks serially and on three kinds of pool with 4 workers."""import osimport sysimport timefrom concurrent.futures import (InterpreterPoolExecutor, ProcessPoolExecutor,ThreadPoolExecutor)from score import score_chunkTASKS, WORKERS = 8, 4def serial() -> int:return sum(score_chunk(seed) for seed in range(TASKS))def pooled(pool_class):def run() -> int:with pool_class(max_workers=WORKERS) as pool:return sum(pool.map(score_chunk, range(TASKS)))return rundef timed(label: str, run) -> None:start = time.perf_counter()result = run()print(f"{label:<13}{time.perf_counter() - start:6.2f} s result {result}")if __name__ == "__main__": # required: worker processes import this modulegil = "on" if sys._is_gil_enabled() else "off"print(f"Python {sys.version.split()[0]}, GIL {gil}, {os.cpu_count()} CPUs")timed("serial", serial)timed("threads", pooled(ThreadPoolExecutor))timed("processes", pooled(ProcessPoolExecutor))timed("interpreters", pooled(InterpreterPoolExecutor))
Read the default build first. Threads were no faster than serial (2.94 s against 2.99 s): with the GIL, pure-Python bytecode from 4 threads takes turns. Processes (0.93 s) and interpreters (0.91 s) were 3.2 and 3.3 times faster, because each has its own GIL and runs on its own CPU. On the free-threaded build, threads were 2.8 times faster than serial (1.28 s against 3.58 s), about as fast as processes, with no pickling of arguments and results and no second copy of the program in memory.
The free-threaded build paid for that in single-threaded speed: 3.58 s serial against 3.09 s for the upstream default build compiled from the same source with the same options, about 16% slower on this loop. The Python documentation reports an average single-threaded overhead of about 1% to 8% on the pyperformance suite; your workload decides where you land.
Interpreters are not free either. Each worker imports its own copy of every module, arguments and results are pickled between interpreters, and an extension module that does not support subinterpreters refuses to load:
"""What an interpreter pool does not share with the program that started it."""from concurrent.futures import InterpreterPoolExecutorseen: list[int] = []def remember(x: int) -> int:seen.append(x) # the list in the worker's interpreter, not oursreturn len(seen)def load(module: str) -> str:__import__(module)return f"{module}: imported"if __name__ == "__main__":with InterpreterPoolExecutor(max_workers=2) as pool:print("worker list sizes:", list(pool.map(remember, range(4))), "| ours:", seen)for name in ("sqlite3", "readline"):try:print(pool.submit(load, name).result())except ImportError as err:print(f"{name}: ImportError: {err}")
Each worker appended to its own seen list, and the program's list stayed empty: module state is per interpreter. sqlite3 loaded, readline did not. Test every third-party C extension you plan to use in an interpreter pool before you rely on it.
C code that releases the GIL
"CPU-bound means processes" is too simple. Much of the heavy work in automation happens in C code that releases the GIL while it runs, and then plain threads use every core even on the default build. hashlib is the common case: it releases the GIL while it hashes more than 2047 bytes in one call.
"""hashlib releases the GIL while it hashes a large buffer, so threads overlap even with the GIL."""import hashlibimport sysimport timefrom concurrent.futures import ThreadPoolExecutorBIG = bytes(256 * 1024 * 1024) # one 256 MiB buffer, hashed once per taskSMALL = [bytes(64)] * 400_000 # 400,000 hashes of 64 bytes per taskdef hash_big(_: int) -> str:return hashlib.sha256(BIG).hexdigest()def hash_small(_: int) -> int:for message in SMALL: # under 2048 bytes, each call keeps the GILhashlib.sha256(message).digest()return len(SMALL)def timed(label: str, fn, workers: int) -> None:start = time.perf_counter()with ThreadPoolExecutor(max_workers=workers) as pool:list(pool.map(fn, range(8)))print(f"{label:<24}{time.perf_counter() - start:6.2f} s")if __name__ == "__main__":print(f"GIL {'on' if sys._is_gil_enabled() else 'off'}, 8 tasks")timed("256 MiB, 1 thread", hash_big, 1)timed("256 MiB, 4 threads", hash_big, 4)timed("64 bytes, 1 thread", hash_small, 1)timed("64 bytes, 4 threads", hash_small, 4)
On the default build, 4 threads hashed 2 GiB in 0.36 s against 0.75 s for one thread, twice as fast (the lab's run on upstream 3.14.7 measured 3.8 times). The same threads gained nothing on 64-byte messages, where the GIL is kept for each call and the loop around them is Python bytecode. For "hash thousands of files" this means hashlib.file_digest (from "Hashes, HMAC and password storage" in py-sec) in a thread pool: the reads release the GIL too, and a process pool would only add the cost of sending data between processes.
asyncio basics: coroutines, await and the event loop
asyncio runs many waiting jobs in one thread. A function defined with async def is a coroutine function: calling it runs none of its body and returns a coroutine object, a piece of work that has not started. await runs an awaitable to its result; when that has to wait (for a socket, a timer), await suspends this coroutine and hands control back to the event loop until the thing is ready. The event loop is a scheduler in one thread: it runs one ready task until its next await, then the next ready one. asyncio.run(main()) creates the loop, runs main() to the end and closes the loop; it is the entry point of an asyncio program.
A coroutine you only await runs on its own. To run several at once, make them tasks: tg.create_task() inside async with asyncio.TaskGroup() as tg: starts each one and supervises them all. The block ends when every task has finished, and if one fails the group cancels the rest. asyncio.timeout(s) cancels the work inside it when the time is up and raises TimeoutError. asyncio.Semaphore(n) lets at most n tasks into an async with block at a time:
"""asyncio basics: a coroutine, await, a task group, a timeout and a semaphore."""import asyncioimport timeasync def check(host: str, seconds: float) -> str:await asyncio.sleep(seconds) # stands in for a network wait; the loop runs other tasks meanwhilereturn f"{host} ok"async def limited(gate: asyncio.Semaphore, host: str) -> str:async with gate: # waits here while two other tasks hold the semaphorereturn await check(host, 0.3)async def main() -> None:coro = check("web01", 0.3)print("check() returned a", type(coro).__name__, "and nothing has run yet")t = time.perf_counter()print(await coro, f"after {time.perf_counter() - t:.1f} s")t = time.perf_counter()async with asyncio.TaskGroup() as tg: # start the tasks, then wait for all of themtasks = [tg.create_task(check(host, 0.3)) for host in ("web02", "web03", "web04")]print([task.result() for task in tasks], f"after {time.perf_counter() - t:.1f} s")t = time.perf_counter()try:async with asyncio.timeout(0.1): # cancels the await inside when the time is upawait check("web05", 0.3)except TimeoutError:print(f"web05 timed out after {time.perf_counter() - t:.1f} s")t = time.perf_counter()gate = asyncio.Semaphore(2)async with asyncio.TaskGroup() as tg:for i in range(4):tg.create_task(limited(gate, f"db{i}"))print(f"4 checks, at most 2 at a time: {time.perf_counter() - t:.1f} s")if __name__ == "__main__":asyncio.run(main()) # creates the event loop, runs main() to the end, closes the loop
Calling check() returned a coroutine and did nothing (one that is never awaited only earns a RuntimeWarning). Awaiting it took 0.3 s. The three tasks in the group waited at the same time, 0.3 s in total instead of 0.9 s. The timeout cancelled web05 at 0.1 s. With the semaphore at 2, four 0.3 s checks ran in two rounds, 0.6 s. The loop switched tasks only at an await, which is the rule the next section tests.
I/O-bound fan-out: threads or asyncio
Waiting on the network releases the GIL, so for I/O the question is how many things wait at once, not which build you run. The lab API answers every request after 0.2 seconds:
"""Lab HTTP service on 127.0.0.1:18509: every GET waits 0.2 s, then answers 200 (a slow API)."""import sysimport timefrom http.server import BaseHTTPRequestHandler, ThreadingHTTPServerclass Handler(BaseHTTPRequestHandler):def do_GET(self):time.sleep(0.2)body = b"ok\n"self.send_response(200)self.send_header("Content-Length", str(len(body)))self.end_headers()self.wfile.write(body)def log_message(self, *args): # one line per request would drown the lab outputpassclass Server(ThreadingHTTPServer):request_queue_size = 512 # the default backlog of 5 would drop connections in a burstdaemon_threads = Trueif __name__ == "__main__":server = Server(("127.0.0.1", 18509), Handler) # binds the port before the message belowprint("slowserver: listening on http://127.0.0.1:18509", file=sys.stderr, flush=True)server.serve_forever()
"""200 GET requests to the slow lab API: a thread pool versus asyncio tasks."""import asyncioimport timeimport urllib.requestfrom concurrent.futures import ThreadPoolExecutorHOST, PORT, REQUESTS = "127.0.0.1", 18509, 200URL = f"http://{HOST}:{PORT}/item"def get_blocking(i: int) -> int:with urllib.request.urlopen(f"{URL}/{i}", timeout=5) as resp:return resp.statusdef with_threads(workers: int) -> list[int]:with ThreadPoolExecutor(max_workers=workers) as pool:return list(pool.map(get_blocking, range(REQUESTS)))async def get_async(i: int) -> int:reader, writer = await asyncio.open_connection(HOST, PORT)try:writer.write(f"GET /item/{i} HTTP/1.0\r\nHost: {HOST}\r\n\r\n".encode())status_line = await reader.readline() # b"HTTP/1.0 200 OK\r\n"await reader.read() # HTTP/1.0: the server closes when donereturn int(status_line.split()[1])finally:writer.close()await writer.wait_closed()async def with_asyncio() -> list[int]:async with asyncio.timeout(5): # one deadline for the whole batchasync with asyncio.TaskGroup() as tg:tasks = [tg.create_task(get_async(i)) for i in range(REQUESTS)]return [t.result() for t in tasks]def timed(label: str, run) -> None:start = time.perf_counter()statuses = run()print(f"{label:<18}{time.perf_counter() - start:6.2f} s {statuses.count(200)} x 200")if __name__ == "__main__":print(f"{REQUESTS} requests, each answered after 0.2 s")timed("threads, 20", lambda: with_threads(20))timed("threads, 200", lambda: with_threads(200))timed("asyncio tasks", lambda: asyncio.run(with_asyncio()))
asyncio.open_connection() returns a reader and a writer for one TCP connection; each await on them is a point where the loop runs other tasks. Start the server in the background, run the fan-out, then stop the server:
The first command saves the server's PID in server.pid and waits at most 10 s for the "listening" line, printed only once the port is bound. The last one stops it: 0 listening sockets on port 18509 means it is gone.
Sequentially, 200 requests would take 40 seconds. 20 threads took 2.09 s, 10 rounds of 0.2 s, because only 20 requests can wait at a time. 200 threads took 0.33 s and 200 asyncio tasks 0.24 s. At this size threads are fine. asyncio earns its place when you have thousands of connections (a task costs far less memory than a thread), and when you need one deadline over the whole batch: asyncio.timeout(5) cancels every task still running, and asyncio.TaskGroup cancels the rest when one task fails, then raises the failures as an ExceptionGroup. A thread cannot be cancelled from outside. asyncio has one rule that decides whether any of this works: nothing may block the event loop.
"""A heartbeat task that should tick every 0.1 s, next to a 1.5 s call, inside asyncio.timeout(0.5)."""import asyncioimport sysimport timeasync def heartbeat(gaps: list[float]) -> None:last = time.monotonic()while True:await asyncio.sleep(0.1)now = time.monotonic()gaps.append(now - last)last = nowasync def slow_call(mode: str) -> None:if mode == "blocking":time.sleep(1.5) # blocks the whole event loopelse:await asyncio.to_thread(time.sleep, 1.5) # runs in a worker thread; the loop keeps goingasync def main(mode: str) -> None:gaps: list[float] = []beat = asyncio.create_task(heartbeat(gaps))await asyncio.sleep(0.3) # let the heartbeat tick a few times firststart = time.monotonic()try:async with asyncio.timeout(0.5):await slow_call(mode)print(f"call finished after {time.monotonic() - start:.2f} s, no timeout")except TimeoutError:print(f"timeout after {time.monotonic() - start:.2f} s")await asyncio.sleep(0.3)beat.cancel()print(f"longest heartbeat gap: {max(gaps, default=0):.2f} s")if __name__ == "__main__":t0 = time.monotonic()asyncio.run(main(sys.argv[1]))print(f"program exited after {time.monotonic() - t0:.2f} s")
With time.sleep(1.5) called directly in a coroutine, the heartbeat that should tick every 0.1 s waited 1.6 s, and the 0.5 s timeout never fired: asyncio.timeout cancels a task at its next await, and blocking code has none. asyncio.to_thread() moved the call to a worker thread. The heartbeat kept ticking and the timeout fired at 0.5 s, but the program still exited at 1.8 s: the timeout stopped the waiting, not the thread, and asyncio.run() waits for its worker threads before it returns. Give the blocking call its own timeout as well (every network call should have one).
Shared state: a counter without a lock
"""Four threads add 1 to a shared counter 250,000 times each: without and with a lock."""import sysimport threadingTHREADS, ADDS = 4, 250_000class Counter:def __init__(self) -> None:self.value = 0self.lock = threading.Lock()def plus_one(n: int) -> int:return n + 1def add_unlocked(c: Counter) -> None:for _ in range(ADDS):c.value += 1 # read, add, write: another thread can run in betweendef add_unlocked_call(c: Counter) -> None:for _ in range(ADDS):c.value = plus_one(c.value) # the same, with a function call between read and writedef add_locked(c: Counter) -> None:for _ in range(ADDS):with c.lock:c.value += 1def run(worker) -> int:c = Counter()threads = [threading.Thread(target=worker, args=(c,)) for _ in range(THREADS)]for t in threads:t.start()for t in threads:t.join()return c.valueif __name__ == "__main__":unlocked = add_unlocked_call if "--call" in sys.argv else add_unlockedprint(f"GIL {'on' if sys._is_gil_enabled() else 'off'}, expected {THREADS * ADDS}:"f" no lock {run(unlocked)}, with lock {run(add_locked)}")
c.value += 1 is three steps: read, add, write back. On the free-threaded build the four threads interleaved those steps and about two thirds of the updates were lost. On the default build the plain += came out right, but only because of where CPython switches threads: it checks for a switch at function calls and loop jumps, not between the read and the write of this line. Put a call between them (--call) and the default build lost updates in one of five runs here; how many varies from run to run. The lock gave the right total on both builds every time. Treat a correct total on the GIL build as luck, not as proof: shared mutable state needs a lock, a queue, or a design that does not share it, whichever build runs your code. The free-threaded build keeps dict, list and set internally consistent, but a read-modify-write across two operations is still yours to protect.
Try this
Start slowserver.py again with the command above (and stop it the same way when you are done). Copy fanout.py to fanout_limit.py and give the asyncio path an asyncio.Semaphore so that at most N requests are in flight. Predict the time for --limit 20 (200 requests, 20 at a time, 0.2 s each), then run it: the lab measured 2.10 s. Then break it: make a coroutine call get_blocking() directly for 10 requests and predict again; the lab measured 2.09 s, the same as sequential, although a TaskGroup started all 10 tasks. Fix it with asyncio.to_thread(get_blocking, i) and explain the 0.42 s you get: the default thread pool on this 4-CPU VM has 8 workers (min(32, CPUs + 4)), so 10 requests take two rounds.
Takeaway
Know which build runs your code, then pick by where the time goes: processes or an interpreter pool for pure-Python CPU work on the GIL build, threads for C code that releases the GIL and for moderate I/O, asyncio for large fan-out with one deadline, and a lock for every read-modify-write on shared state on either build.
python3, a ThreadPoolExecutor(8) gives no speed-up on an 8-core host. What is the best next step?__main__ guard.async with asyncio.timeout(1):. Health checks on the same event loop start failing, and the timeout never fires. Which change fixes both problems?