Concurrency on Python 3.14: threads, processes, interpreters and asyncio

The GIL build and the free-threaded build, forkserver, and asyncio task groups.

Expert45 min · lesson 9 of 15
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (11 files, 4 KB): scr-py-concurrency.tar.gz. Unpack it with tar -xzf scr-py-concurrency.tar.gz, which creates scr-py-concurrency/. SHA-256: 9c00d27e717ce040abdb828ea395d6e8beed0fdf25b6ea80a1d7761cc57a3ca8

This lesson gives you measured answers to "should this run in threads, processes, interpreters or asyncio?" on Python 3.14. You will check which build of CPython you are running, time the same CPU-bound job four ways on the default build and on the free-threaded build, see C code that releases the GIL, learn how asyncio runs coroutines, fan out 200 slow HTTP calls with threads and with asyncio, stall an event loop on purpose, and watch a shared counter lose updates. Every number below was measured on the lab VM (4 vCPUs, Ubuntu 26.04). Timings change from run to run and from machine to machine; read them as ratios, and measure your own workload before you decide.

Refresher: "Functions, modules and exceptions" in py-sec owns if __name__ == "__main__":, and "Failure design" (the previous lesson) owns ExceptionGroup and except*. This lesson owns asyncio; the fleet lesson and the capstone build on it. Unpack the lesson files (the box at the top of this page) in your home directory and work in ~/scr-py-concurrency: they hold every script below, including slowserver.py, the small HTTP server the fan-out section starts and stops.

Two builds of CPython 3.14

The global interpreter lock (GIL) is a lock inside CPython that lets only one thread at a time run Python bytecode in one interpreter. It is a property of a build, not of the language. Python 3.14 comes in two builds: the default build, which keeps the GIL, and the free-threaded build, compiled with --disable-gil, where threads run Python code in parallel. PEP 703 introduced the free-threaded build in 3.13 as experimental; PEP 779 made it officially supported in 3.14, but it stays an optional, separate build (its binary is python3.14t), and nothing has been decided about making it the default. Ask the interpreter which one you have:

which_python.py
import sys
import sysconfig
print(sys.version)
# Py_GIL_DISABLED is fixed when the interpreter is compiled; _is_gil_enabled() is the state now.
print("free-threaded build:", sysconfig.get_config_var("Py_GIL_DISABLED") == 1)
print("GIL enabled now: ", sys._is_gil_enabled())
deploy@web01:~/scr-py-concurrency · Ubuntu 26.04 LTS
$ python3 which_python.py
3.14.4 (main, Aug 20 2026, 10:41:58) [GCC 15.2.0] free-threaded build: False GIL enabled now: True
$ /opt/cpython-3.14.7t/bin/python3 which_python.py
3.14.7 free-threading build (main, Sep 27 2026, 14:24:43) [GCC 15.2.0] free-threaded build: True GIL enabled now: False
$ PYTHON_GIL=1 /opt/cpython-3.14.7t/bin/python3 which_python.py
3.14.7 free-threading build (main, Sep 27 2026, 14:24:43) [GCC 15.2.0] free-threaded build: True GIL enabled now: True
$ PYTHON_GIL=0 python3 which_python.py
Fatal Python error: config_read_gil: Disabling the GIL is not supported by this build Python runtime state: preinitialized

Ubuntu's python3 is the default build: Py_GIL_DISABLED is not set and the GIL is on. /opt/cpython-3.14.7t is the lab's free-threaded build of the upstream 3.14.7 source; its version string says "free-threading build" and the GIL is off. Py_GIL_DISABLED is fixed when the interpreter is compiled; sys._is_gil_enabled() is the state right now, and PYTHON_GIL=1 (or -X gil=1) turns the GIL back on in a free-threaded build. A C extension that is not marked as safe without the GIL also turns it back on when imported, with a warning, so check at run time, not only at install time. The default build cannot go the other way: PYTHON_GIL=0 is a fatal error there.

Ubuntu 26.04 packages neither the free-threaded build nor 3.14.7. The archive has python3.14 at 3.14.4, and no package whose name looks like a free-threaded Python:

deploy@web01:~/scr-py-concurrency · Ubuntu 26.04 LTS
$ apt-cache policy python3.14 | head -3
python3.14: Installed: 3.14.4-1ubuntu0.2 Candidate: 3.14.4-1ubuntu0.2
$ apt-cache search --names-only 'python3\.14t|nogil|freethreading|free-threaded' | wc -l
0

So the lab VM built upstream 3.14.7 from the python.org source twice when it was prepared: once with --disable-gil into /opt/cpython-3.14.7t, once without into /opt/cpython-3.14.7. These are the steps for the free-threaded one. A build takes several minutes on 4 CPUs, so the lab does not repeat it; the next terminal proves what was built:

build-3.14t.sh (what the lab VM ran once, when it was prepared)
# Build dependencies from the Ubuntu archive
sudo apt-get install -y build-essential pkg-config libssl-dev zlib1g-dev libbz2-dev \
libreadline-dev libsqlite3-dev libffi-dev liblzma-dev uuid-dev libgdbm-dev \
libgdbm-compat-dev libzstd-dev
# The source, checked against the SHA-256 that python.org publishes for this release
curl -fsSLO https://www.python.org/ftp/python/3.14.7/Python-3.14.7.tar.xz
echo '3b48dac8fb59f62eaa67ac83c1eb12bda1b7a08406dd286e252c11a66be27f81 Python-3.14.7.tar.xz' | sha256sum -c
tar -xf Python-3.14.7.tar.xz && cd Python-3.14.7
# Its own prefix: /usr/bin/python3 and everything Ubuntu installed with it stay untouched
./configure --prefix=/opt/cpython-3.14.7t --disable-gil
make -j"$(nproc)"
sudo make install
sudo ln -sf python3.14t /opt/cpython-3.14.7t/bin/python3 # the same python3 name inside the prefix
# The default build for comparison: a fresh copy of the source, ./configure --prefix=/opt/cpython-3.14.7

CONFIG_ARGS records the options an interpreter was configured with. Check all three:

deploy@web01:~/scr-py-concurrency · Ubuntu 26.04 LTS
$ for py in python3 /opt/cpython-3.14.7t/bin/python3 /opt/cpython-3.14.7/bin/python3; do "$py" -VV "$py" -c "import sysconfig; print(sysconfig.get_config_var('CONFIG_ARGS'))" done
Python 3.14.4 (main, Aug 20 2026, 10:41:58) [GCC 15.2.0] '--enable-shared' '--prefix=/usr' '--libdir=/usr/lib/aarch64-linux-gnu' '--enable-ipv6' '--enable-loadable-sqlite-extensions' '--with-dbmliborder=bdb:gdbm' '--with-computed-gotos' '--without-ensurepip' '--with-system-expat' '--with-dtrace' '--with-ssl-default-suites=openssl' '--with-system-libmpdec=no' '--with-wheel-pkg-dir=/usr/share/python-wheels/' 'MKDIR_P=/bin/mkdir -p' '--enable-experimental-jit=yes-off' 'CC=aarch64-linux-gnu-gcc' Python 3.14.7 free-threading build (main, Sep 27 2026, 14:24:43) [GCC 15.2.0] '--prefix=/opt/cpython-3.14.7t' '--disable-gil' Python 3.14.7 (main, Sep 27 2026, 14:23:24) [GCC 15.2.0] '--prefix=/opt/cpython-3.14.7'

The two /opt builds differ only in --disable-gil, which makes the timing comparison below fair. Ubuntu's own build does not use it either. Other routes to a free-threaded build: the python.org installers for macOS and Windows offer it as an optional component, and uv python install 3.14t downloads a prebuilt one (uv arrives in "Shipping automation" later in this course). If you have no free-threaded build, read the terminals that call /opt/cpython-3.14.7t; every other command runs on Ubuntu's python3.

CPU-bound Python: threads, processes and interpreters

The job is pure Python: generate host names and flag the random-looking ones by their character entropy, the way a DGA detector scores domains. score_chunk(seed) does 150,000 names and shares nothing with other calls:

score.py
"""CPU-bound work in pure Python: flag generated host names that look random (DGA-like)."""
import math
import random
LETTERS = "abcdefghijklmnopqrstuvwxyz0123456789"
def entropy(name: str) -> float:
counts: dict[str, int] = {}
for ch in name:
counts[ch] = counts.get(ch, 0) + 1
n = len(name)
return -sum(c / n * math.log2(c / n) for c in counts.values())
def score_chunk(seed: int, count: int = 150_000) -> int:
"""Generate count names from seed; return how many score above the threshold."""
rng = random.Random(seed) # one generator per task: nothing is shared between workers
flagged = 0
for _ in range(count):
name = "".join(rng.choices(LETTERS, k=16))
if entropy(name) > 3.6:
flagged += 1
return flagged

compare.py runs 8 of those tasks serially, then on a pool of 4 threads, 4 processes and 4 interpreters. InterpreterPoolExecutor is new in 3.14: each worker is a separate interpreter inside the same process, with its own GIL. Before running it, see what the 3.14 default for processes on Linux does to a script without a __main__ guard:

no_guard.py
from concurrent.futures import ProcessPoolExecutor
from score import score_chunk
# No __main__ guard: the fork server imports this file, so it would start a pool too.
with ProcessPoolExecutor(max_workers=2) as pool:
print(sum(pool.map(score_chunk, range(4))))
deploy@web01:~/scr-py-concurrency · Ubuntu 26.04 LTS
$ python3 -c "import multiprocessing as mp; print(mp.get_start_method())"
forkserver
$ python3 no_guard.py
Traceback (most recent call last): … File "/usr/lib/python3.14/multiprocessing/forkserver.py", line 230, in main spawn.import_main_path(main_path) … File "/home/deploy/scr-py-concurrency/no_guard.py", line 7, in <module> print(sum(pool.map(score_chunk, range(4)))) … RuntimeError: An attempt has been made to start a new process before the current process has finished its bootstrapping phase. … Traceback (most recent call last): File "/home/deploy/scr-py-concurrency/no_guard.py", line 7, in <module> print(sum(pool.map(score_chunk, range(4)))) … ConnectionResetError: [Errno 104] Connection reset by peer

Since 3.14 the default start method on Linux is forkserver, not fork. With fork, each worker was a copy of the running program. With forkserver, the pool first starts a server process, which imports your main module once (spawn.import_main_path in the first traceback) so that the workers it forks from itself start with your functions loaded. Without the guard, your top-level code runs again inside that server and tries to start a pool of its own; multiprocessing stops it with the "bootstrapping phase" RuntimeError, and the server dies. The second traceback is the parent: it was talking to the server when the server exited, hence ConnectionResetError. (With spawn, the default on macOS and Windows, every worker imports the main module, so the guard is needed there too.) The fix is the guard, and the function a worker runs must be importable from a module, which is why score_chunk lives in score.py:

compare.py
"""Run the same 8 CPU-bound tasks serially and on three kinds of pool with 4 workers."""
import os
import sys
import time
from concurrent.futures import (InterpreterPoolExecutor, ProcessPoolExecutor,
ThreadPoolExecutor)
from score import score_chunk
TASKS, WORKERS = 8, 4
def serial() -> int:
return sum(score_chunk(seed) for seed in range(TASKS))
def pooled(pool_class):
def run() -> int:
with pool_class(max_workers=WORKERS) as pool:
return sum(pool.map(score_chunk, range(TASKS)))
return run
def timed(label: str, run) -> None:
start = time.perf_counter()
result = run()
print(f"{label:<13}{time.perf_counter() - start:6.2f} s result {result}")
if __name__ == "__main__": # required: worker processes import this module
gil = "on" if sys._is_gil_enabled() else "off"
print(f"Python {sys.version.split()[0]}, GIL {gil}, {os.cpu_count()} CPUs")
timed("serial", serial)
timed("threads", pooled(ThreadPoolExecutor))
timed("processes", pooled(ProcessPoolExecutor))
timed("interpreters", pooled(InterpreterPoolExecutor))
deploy@web01:~/scr-py-concurrency · Ubuntu 26.04 LTS
$ python3 compare.py
Python 3.14.4, GIL on, 4 CPUs serial 2.99 s result 715344 threads 2.94 s result 715344 processes 0.93 s result 715344 interpreters 0.91 s result 715344
$ /opt/cpython-3.14.7t/bin/python3 compare.py
Python 3.14.7, GIL off, 4 CPUs serial 3.58 s result 715344 threads 1.28 s result 715344 processes 1.24 s result 715344 interpreters 1.13 s result 715344
$ /opt/cpython-3.14.7/bin/python3 compare.py
Python 3.14.7, GIL on, 4 CPUs serial 3.09 s result 715344 …

Read the default build first. Threads were no faster than serial (2.94 s against 2.99 s): with the GIL, pure-Python bytecode from 4 threads takes turns. Processes (0.93 s) and interpreters (0.91 s) were 3.2 and 3.3 times faster, because each has its own GIL and runs on its own CPU. On the free-threaded build, threads were 2.8 times faster than serial (1.28 s against 3.58 s), about as fast as processes, with no pickling of arguments and results and no second copy of the program in memory.

The free-threaded build paid for that in single-threaded speed: 3.58 s serial against 3.09 s for the upstream default build compiled from the same source with the same options, about 16% slower on this loop. The Python documentation reports an average single-threaded overhead of about 1% to 8% on the pyperformance suite; your workload decides where you land.

Interpreters are not free either. Each worker imports its own copy of every module, arguments and results are pickled between interpreters, and an extension module that does not support subinterpreters refuses to load:

interp_limits.py
"""What an interpreter pool does not share with the program that started it."""
from concurrent.futures import InterpreterPoolExecutor
seen: list[int] = []
def remember(x: int) -> int:
seen.append(x) # the list in the worker's interpreter, not ours
return len(seen)
def load(module: str) -> str:
__import__(module)
return f"{module}: imported"
if __name__ == "__main__":
with InterpreterPoolExecutor(max_workers=2) as pool:
print("worker list sizes:", list(pool.map(remember, range(4))), "| ours:", seen)
for name in ("sqlite3", "readline"):
try:
print(pool.submit(load, name).result())
except ImportError as err:
print(f"{name}: ImportError: {err}")
deploy@web01:~/scr-py-concurrency · Ubuntu 26.04 LTS
$ python3 interp_limits.py
worker list sizes: [1, 1, 2, 2] | ours: [] sqlite3: imported readline: ImportError: module readline does not support loading in subinterpreters

Each worker appended to its own seen list, and the program's list stayed empty: module state is per interpreter. sqlite3 loaded, readline did not. Test every third-party C extension you plan to use in an interpreter pool before you rely on it.

C code that releases the GIL

"CPU-bound means processes" is too simple. Much of the heavy work in automation happens in C code that releases the GIL while it runs, and then plain threads use every core even on the default build. hashlib is the common case: it releases the GIL while it hashes more than 2047 bytes in one call.

hash_threads.py
"""hashlib releases the GIL while it hashes a large buffer, so threads overlap even with the GIL."""
import hashlib
import sys
import time
from concurrent.futures import ThreadPoolExecutor
BIG = bytes(256 * 1024 * 1024) # one 256 MiB buffer, hashed once per task
SMALL = [bytes(64)] * 400_000 # 400,000 hashes of 64 bytes per task
def hash_big(_: int) -> str:
return hashlib.sha256(BIG).hexdigest()
def hash_small(_: int) -> int:
for message in SMALL: # under 2048 bytes, each call keeps the GIL
hashlib.sha256(message).digest()
return len(SMALL)
def timed(label: str, fn, workers: int) -> None:
start = time.perf_counter()
with ThreadPoolExecutor(max_workers=workers) as pool:
list(pool.map(fn, range(8)))
print(f"{label:<24}{time.perf_counter() - start:6.2f} s")
if __name__ == "__main__":
print(f"GIL {'on' if sys._is_gil_enabled() else 'off'}, 8 tasks")
timed("256 MiB, 1 thread", hash_big, 1)
timed("256 MiB, 4 threads", hash_big, 4)
timed("64 bytes, 1 thread", hash_small, 1)
timed("64 bytes, 4 threads", hash_small, 4)
deploy@web01:~/scr-py-concurrency · Ubuntu 26.04 LTS
$ python3 hash_threads.py
GIL on, 8 tasks 256 MiB, 1 thread 0.75 s 256 MiB, 4 threads 0.36 s 64 bytes, 1 thread 0.75 s 64 bytes, 4 threads 0.74 s

On the default build, 4 threads hashed 2 GiB in 0.36 s against 0.75 s for one thread, twice as fast (the lab's run on upstream 3.14.7 measured 3.8 times). The same threads gained nothing on 64-byte messages, where the GIL is kept for each call and the loop around them is Python bytecode. For "hash thousands of files" this means hashlib.file_digest (from "Hashes, HMAC and password storage" in py-sec) in a thread pool: the reads release the GIL too, and a process pool would only add the cost of sending data between processes.

asyncio basics: coroutines, await and the event loop

asyncio runs many waiting jobs in one thread. A function defined with async def is a coroutine function: calling it runs none of its body and returns a coroutine object, a piece of work that has not started. await runs an awaitable to its result; when that has to wait (for a socket, a timer), await suspends this coroutine and hands control back to the event loop until the thing is ready. The event loop is a scheduler in one thread: it runs one ready task until its next await, then the next ready one. asyncio.run(main()) creates the loop, runs main() to the end and closes the loop; it is the entry point of an asyncio program.

A coroutine you only await runs on its own. To run several at once, make them tasks: tg.create_task() inside async with asyncio.TaskGroup() as tg: starts each one and supervises them all. The block ends when every task has finished, and if one fails the group cancels the rest. asyncio.timeout(s) cancels the work inside it when the time is up and raises TimeoutError. asyncio.Semaphore(n) lets at most n tasks into an async with block at a time:

async_basics.py
"""asyncio basics: a coroutine, await, a task group, a timeout and a semaphore."""
import asyncio
import time
async def check(host: str, seconds: float) -> str:
await asyncio.sleep(seconds) # stands in for a network wait; the loop runs other tasks meanwhile
return f"{host} ok"
async def limited(gate: asyncio.Semaphore, host: str) -> str:
async with gate: # waits here while two other tasks hold the semaphore
return await check(host, 0.3)
async def main() -> None:
coro = check("web01", 0.3)
print("check() returned a", type(coro).__name__, "and nothing has run yet")
t = time.perf_counter()
print(await coro, f"after {time.perf_counter() - t:.1f} s")
t = time.perf_counter()
async with asyncio.TaskGroup() as tg: # start the tasks, then wait for all of them
tasks = [tg.create_task(check(host, 0.3)) for host in ("web02", "web03", "web04")]
print([task.result() for task in tasks], f"after {time.perf_counter() - t:.1f} s")
t = time.perf_counter()
try:
async with asyncio.timeout(0.1): # cancels the await inside when the time is up
await check("web05", 0.3)
except TimeoutError:
print(f"web05 timed out after {time.perf_counter() - t:.1f} s")
t = time.perf_counter()
gate = asyncio.Semaphore(2)
async with asyncio.TaskGroup() as tg:
for i in range(4):
tg.create_task(limited(gate, f"db{i}"))
print(f"4 checks, at most 2 at a time: {time.perf_counter() - t:.1f} s")
if __name__ == "__main__":
asyncio.run(main()) # creates the event loop, runs main() to the end, closes the loop
deploy@web01:~/scr-py-concurrency · Ubuntu 26.04 LTS
$ python3 async_basics.py
check() returned a coroutine and nothing has run yet web01 ok after 0.3 s ['web02 ok', 'web03 ok', 'web04 ok'] after 0.3 s web05 timed out after 0.1 s 4 checks, at most 2 at a time: 0.6 s

Calling check() returned a coroutine and did nothing (one that is never awaited only earns a RuntimeWarning). Awaiting it took 0.3 s. The three tasks in the group waited at the same time, 0.3 s in total instead of 0.9 s. The timeout cancelled web05 at 0.1 s. With the semaphore at 2, four 0.3 s checks ran in two rounds, 0.6 s. The loop switched tasks only at an await, which is the rule the next section tests.

I/O-bound fan-out: threads or asyncio

Waiting on the network releases the GIL, so for I/O the question is how many things wait at once, not which build you run. The lab API answers every request after 0.2 seconds:

slowserver.py
"""Lab HTTP service on 127.0.0.1:18509: every GET waits 0.2 s, then answers 200 (a slow API)."""
import sys
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
class Handler(BaseHTTPRequestHandler):
def do_GET(self):
time.sleep(0.2)
body = b"ok\n"
self.send_response(200)
self.send_header("Content-Length", str(len(body)))
self.end_headers()
self.wfile.write(body)
def log_message(self, *args): # one line per request would drown the lab output
pass
class Server(ThreadingHTTPServer):
request_queue_size = 512 # the default backlog of 5 would drop connections in a burst
daemon_threads = True
if __name__ == "__main__":
server = Server(("127.0.0.1", 18509), Handler) # binds the port before the message below
print("slowserver: listening on http://127.0.0.1:18509", file=sys.stderr, flush=True)
server.serve_forever()
fanout.py
"""200 GET requests to the slow lab API: a thread pool versus asyncio tasks."""
import asyncio
import time
import urllib.request
from concurrent.futures import ThreadPoolExecutor
HOST, PORT, REQUESTS = "127.0.0.1", 18509, 200
URL = f"http://{HOST}:{PORT}/item"
def get_blocking(i: int) -> int:
with urllib.request.urlopen(f"{URL}/{i}", timeout=5) as resp:
return resp.status
def with_threads(workers: int) -> list[int]:
with ThreadPoolExecutor(max_workers=workers) as pool:
return list(pool.map(get_blocking, range(REQUESTS)))
async def get_async(i: int) -> int:
reader, writer = await asyncio.open_connection(HOST, PORT)
try:
writer.write(f"GET /item/{i} HTTP/1.0\r\nHost: {HOST}\r\n\r\n".encode())
status_line = await reader.readline() # b"HTTP/1.0 200 OK\r\n"
await reader.read() # HTTP/1.0: the server closes when done
return int(status_line.split()[1])
finally:
writer.close()
await writer.wait_closed()
async def with_asyncio() -> list[int]:
async with asyncio.timeout(5): # one deadline for the whole batch
async with asyncio.TaskGroup() as tg:
tasks = [tg.create_task(get_async(i)) for i in range(REQUESTS)]
return [t.result() for t in tasks]
def timed(label: str, run) -> None:
start = time.perf_counter()
statuses = run()
print(f"{label:<18}{time.perf_counter() - start:6.2f} s {statuses.count(200)} x 200")
if __name__ == "__main__":
print(f"{REQUESTS} requests, each answered after 0.2 s")
timed("threads, 20", lambda: with_threads(20))
timed("threads, 200", lambda: with_threads(200))
timed("asyncio tasks", lambda: asyncio.run(with_asyncio()))

asyncio.open_connection() returns a reader and a writer for one TCP connection; each await on them is a point where the loop runs other tasks. Start the server in the background, run the fan-out, then stop the server:

deploy@web01:~/scr-py-concurrency · Ubuntu 26.04 LTS
$ python3 slowserver.py 2> server.log & echo $! > server.pid timeout 10 bash -c 'until grep -q listening server.log; do sleep 0.2; done'; cat server.log
slowserver: listening on http://127.0.0.1:18509
$ python3 fanout.py
200 requests, each answered after 0.2 s threads, 20 2.09 s 200 x 200 threads, 200 0.33 s 200 x 200 asyncio tasks 0.24 s 200 x 200
$ kill "$(cat server.pid)" && rm server.pid sleep 0.5; ss -Hltn 'sport = :18509' | wc -l
0
# At an interactive prompt Bash also reports the background job as Terminated.

The first command saves the server's PID in server.pid and waits at most 10 s for the "listening" line, printed only once the port is bound. The last one stops it: 0 listening sockets on port 18509 means it is gone.

Sequentially, 200 requests would take 40 seconds. 20 threads took 2.09 s, 10 rounds of 0.2 s, because only 20 requests can wait at a time. 200 threads took 0.33 s and 200 asyncio tasks 0.24 s. At this size threads are fine. asyncio earns its place when you have thousands of connections (a task costs far less memory than a thread), and when you need one deadline over the whole batch: asyncio.timeout(5) cancels every task still running, and asyncio.TaskGroup cancels the rest when one task fails, then raises the failures as an ExceptionGroup. A thread cannot be cancelled from outside. asyncio has one rule that decides whether any of this works: nothing may block the event loop.

loop_stall.py
"""A heartbeat task that should tick every 0.1 s, next to a 1.5 s call, inside asyncio.timeout(0.5)."""
import asyncio
import sys
import time
async def heartbeat(gaps: list[float]) -> None:
last = time.monotonic()
while True:
await asyncio.sleep(0.1)
now = time.monotonic()
gaps.append(now - last)
last = now
async def slow_call(mode: str) -> None:
if mode == "blocking":
time.sleep(1.5) # blocks the whole event loop
else:
await asyncio.to_thread(time.sleep, 1.5) # runs in a worker thread; the loop keeps going
async def main(mode: str) -> None:
gaps: list[float] = []
beat = asyncio.create_task(heartbeat(gaps))
await asyncio.sleep(0.3) # let the heartbeat tick a few times first
start = time.monotonic()
try:
async with asyncio.timeout(0.5):
await slow_call(mode)
print(f"call finished after {time.monotonic() - start:.2f} s, no timeout")
except TimeoutError:
print(f"timeout after {time.monotonic() - start:.2f} s")
await asyncio.sleep(0.3)
beat.cancel()
print(f"longest heartbeat gap: {max(gaps, default=0):.2f} s")
if __name__ == "__main__":
t0 = time.monotonic()
asyncio.run(main(sys.argv[1]))
print(f"program exited after {time.monotonic() - t0:.2f} s")
deploy@web01:~/scr-py-concurrency · Ubuntu 26.04 LTS
$ python3 loop_stall.py blocking
call finished after 1.50 s, no timeout longest heartbeat gap: 1.59 s program exited after 2.11 s
$ python3 loop_stall.py thread
timeout after 0.50 s longest heartbeat gap: 0.11 s program exited after 1.82 s

With time.sleep(1.5) called directly in a coroutine, the heartbeat that should tick every 0.1 s waited 1.6 s, and the 0.5 s timeout never fired: asyncio.timeout cancels a task at its next await, and blocking code has none. asyncio.to_thread() moved the call to a worker thread. The heartbeat kept ticking and the timeout fired at 0.5 s, but the program still exited at 1.8 s: the timeout stopped the waiting, not the thread, and asyncio.run() waits for its worker threads before it returns. Give the blocking call its own timeout as well (every network call should have one).

Shared state: a counter without a lock

race.py
"""Four threads add 1 to a shared counter 250,000 times each: without and with a lock."""
import sys
import threading
THREADS, ADDS = 4, 250_000
class Counter:
def __init__(self) -> None:
self.value = 0
self.lock = threading.Lock()
def plus_one(n: int) -> int:
return n + 1
def add_unlocked(c: Counter) -> None:
for _ in range(ADDS):
c.value += 1 # read, add, write: another thread can run in between
def add_unlocked_call(c: Counter) -> None:
for _ in range(ADDS):
c.value = plus_one(c.value) # the same, with a function call between read and write
def add_locked(c: Counter) -> None:
for _ in range(ADDS):
with c.lock:
c.value += 1
def run(worker) -> int:
c = Counter()
threads = [threading.Thread(target=worker, args=(c,)) for _ in range(THREADS)]
for t in threads:
t.start()
for t in threads:
t.join()
return c.value
if __name__ == "__main__":
unlocked = add_unlocked_call if "--call" in sys.argv else add_unlocked
print(f"GIL {'on' if sys._is_gil_enabled() else 'off'}, expected {THREADS * ADDS}:"
f" no lock {run(unlocked)}, with lock {run(add_locked)}")
deploy@web01:~/scr-py-concurrency · Ubuntu 26.04 LTS
$ /opt/cpython-3.14.7t/bin/python3 race.py
GIL off, expected 1000000: no lock 323535, with lock 1000000
$ python3 race.py
GIL on, expected 1000000: no lock 1000000, with lock 1000000
$ for run in 1 2 3 4 5; do python3 race.py --call; done
GIL on, expected 1000000: no lock 1000000, with lock 1000000 GIL on, expected 1000000: no lock 984072, with lock 1000000 GIL on, expected 1000000: no lock 1000000, with lock 1000000 GIL on, expected 1000000: no lock 1000000, with lock 1000000 GIL on, expected 1000000: no lock 1000000, with lock 1000000

c.value += 1 is three steps: read, add, write back. On the free-threaded build the four threads interleaved those steps and about two thirds of the updates were lost. On the default build the plain += came out right, but only because of where CPython switches threads: it checks for a switch at function calls and loop jumps, not between the read and the write of this line. Put a call between them (--call) and the default build lost updates in one of five runs here; how many varies from run to run. The lock gave the right total on both builds every time. Treat a correct total on the GIL build as luck, not as proof: shared mutable state needs a lock, a queue, or a design that does not share it, whichever build runs your code. The free-threaded build keeps dict, list and set internally consistent, but a read-modify-write across two operations is still yours to protect.

Which model for which job
Pure-Python CPU work
Processes
any build; guard __main__, pickled data
Interpreters
3.14; own GIL each, no shared modules
Threads on 3.14t
parallel, shared memory, needs locks
C code that releases the GIL
Threads
hashlib, zlib, file and socket I/O
Waiting on I/O
Threads
tens to hundreds at a time
asyncio
thousands; one deadline, cancellation
Measure your own workload on your own build; these are starting points, not rules.

Try this

Start slowserver.py again with the command above (and stop it the same way when you are done). Copy fanout.py to fanout_limit.py and give the asyncio path an asyncio.Semaphore so that at most N requests are in flight. Predict the time for --limit 20 (200 requests, 20 at a time, 0.2 s each), then run it: the lab measured 2.10 s. Then break it: make a coroutine call get_blocking() directly for 10 requests and predict again; the lab measured 2.09 s, the same as sequential, although a TaskGroup started all 10 tasks. Fix it with asyncio.to_thread(get_blocking, i) and explain the 0.42 s you get: the default thread pool on this 4-CPU VM has 8 workers (min(32, CPUs + 4)), so 10 requests take two rounds.

Takeaway

Know which build runs your code, then pick by where the time goes: processes or an interpreter pool for pure-Python CPU work on the GIL build, threads for C code that releases the GIL and for moderate I/O, asyncio for large fan-out with one deadline, and a lock for every read-modify-write on shared state on either build.

Quick check
01A nightly job computes risk scores in pure Python over 2 million records. On Ubuntu 26.04's python3, a ThreadPoolExecutor(8) gives no speed-up on an 8-core host. What is the best next step?
Incorrect — On the default build pure-Python bytecode from many threads still takes turns on one GIL; more threads add switching, not speed.
Correct — Each process or interpreter has its own GIL, so pure-Python work runs on several cores; measure, because startup and pickling cost something.
Incorrect — asyncio runs its tasks in one thread; it overlaps waiting, not computing, so CPU-bound scoring gains nothing.
Incorrect — hashlib releases the GIL only while it hashes large buffers; it does not make unrelated Python code run in parallel.
02A script that worked on Python 3.12 now fails on 3.14 on Linux with "An attempt has been made to start a new process before the current process has finished its bootstrapping phase". What changed?
Correct — With fork the child inherited the running program; with forkserver a server process imports the main module first, so top-level code needs the __main__ guard.
Incorrect — ProcessPoolExecutor is still there and ran in this lesson; InterpreterPoolExecutor was added next to it, not instead of it.
Incorrect — The default build still has the GIL (PEP 779 keeps the free-threaded build optional), and both builds can start processes.
Incorrect — Nothing in the kernel changed; Python chose a different default start method, and the fix is in your script.
03An asyncio service wraps a vendor SDK call (blocking, about 3 s) in async with asyncio.timeout(1):. Health checks on the same event loop start failing, and the timeout never fires. Which change fixes both problems?
Incorrect — A timeout cancels a task at its next await; a blocking call has none, so no value makes it fire during the call.
Incorrect — wait_for uses the same cancellation mechanism; it cannot interrupt code that never returns control to the loop.
Incorrect — The event loop still runs in one thread; a blocking call in a coroutine stops it on any build.
Correct — The loop keeps running and the asyncio timeout can fire, but it cannot stop the thread, so the call needs its own timeout.

Related