asyncio, threads or processes? Choosing Python concurrency by workload
A practical decision guide: which Python concurrency model fits I/O-bound, CPU-bound and mixed workloads, with small runnable examples and the traps to avoid.
Python gives you three ways to do more than one thing at a time: asyncio, threads and processes. They are not interchangeable. Picking the wrong one either gives you no speed-up at all or a lot of complexity for very little gain. The right choice depends almost entirely on what your program is waiting for.
Start with one question: is the work I/O-bound or CPU-bound?
- I/O-bound work spends most of its time waiting: HTTP calls, database queries, reading files, talking to an LLM API. The CPU is idle while you wait.
- CPU-bound work keeps a core busy: parsing large files, image processing, numeric loops in pure Python, compressing data.
In CPython, the Global Interpreter Lock (GIL) lets only one thread run Python bytecode at a time. That single fact drives the whole decision. (Free-threaded builds of CPython exist, but most production deployments and libraries still assume the GIL, so plan for it.)
| Workload | Best fit | Why |
|---|---|---|
| Many network calls, one service | asyncio |
Thousands of waits on one thread, very low overhead |
| Network calls with a blocking library | threads | Blocking calls release the GIL while they wait |
| Heavy pure-Python computation | processes | Each process has its own interpreter and GIL |
| Mixed | asyncio + a process pool |
Keep the event loop free, push CPU work out |
I/O-bound: asyncio when your libraries support it
If your HTTP client, database driver and SDKs are async-native, asyncio is the most efficient option. One thread can keep hundreds of requests in flight.
import asyncio
import httpx
async def fetch_all(urls):
async with httpx.AsyncClient(timeout=10) as client:
tasks = [client.get(u) for u in urls]
return await asyncio.gather(*tasks, return_exceptions=True)
results = asyncio.run(fetch_all(["https://example.com"] * 50))
Two rules keep async code healthy:
- Never call blocking code inside a coroutine. A single
time.sleep()or synchronousrequests.get()freezes every other task on the loop. Useasyncio.to_thread(func, ...)for the occasional blocking call. - Bound your concurrency. Firing 10,000 requests at once will hit rate limits or exhaust connections. A semaphore fixes it:
sem = asyncio.Semaphore(20)
async def limited(client, url):
async with sem:
return await client.get(url)
I/O-bound with blocking libraries: threads
Plenty of good libraries are synchronous. Threads work well here because the GIL is released while a thread waits on a socket or a file.
from concurrent.futures import ThreadPoolExecutor
import requests
def get(url):
return requests.get(url, timeout=10).status_code
with ThreadPoolExecutor(max_workers=16) as pool:
codes = list(pool.map(get, urls))
Keep thread counts modest (tens, not thousands) and treat any shared mutable state with care: use a queue.Queue or a lock rather than appending to a shared list from many threads.
CPU-bound: processes
For pure-Python computation, threads won't help, because only one of them runs at a time. Use a process pool so each worker gets its own interpreter:
from concurrent.futures import ProcessPoolExecutor
def score(chunk):
return sum(len(line.split()) for line in chunk)
if __name__ == "__main__":
chunks = [lines[i:i + 10_000] for i in range(0, len(lines), 10_000)]
with ProcessPoolExecutor() as pool:
total = sum(pool.map(score, chunks))
Things to know:
- Arguments and results are pickled between processes. Send work in reasonably large chunks, or the copying costs more than the computation.
- Always guard the entry point with
if __name__ == "__main__":. On Windows and macOS, processes are started by re-importing your module. - Before reaching for processes, check whether a vectorised library (NumPy, Polars, DuckDB) already does the heavy lifting in native code. That is often a bigger win.
Mixed workloads: keep the event loop free
A web service built on asyncio sometimes has to do something CPU-heavy, like generating a report. Run it in a process pool so request handling keeps flowing:
import asyncio
from concurrent.futures import ProcessPoolExecutor
pool = ProcessPoolExecutor(max_workers=2)
async def handler(data):
loop = asyncio.get_running_loop()
return await loop.run_in_executor(pool, build_report, data)
A quick checklist
- Measure first. Profile with
cProfileorpy-spyto see whether you are waiting or computing. - Waiting on the network and your libraries are async? Use asyncio with a semaphore.
- Waiting, but your libraries are blocking? Use a thread pool.
- Computing in pure Python? Use a process pool, or move the hot loop into a native library.
- Mixing both? asyncio for the I/O,
run_in_executorwith a process pool for the CPU work.
The best concurrency model is the simplest one that removes your actual bottleneck. Often that is a thread pool with eight workers, not a rewrite.
Written by RecallRun Editors for the RecallRun community. Community posts are checked for safety and reviewed by our editors before publishing, but the views and claims are the author's own. Links are the author's; open them with care. Report this post.
More from the community
- Tech articles
Logs, metrics and traces: a practical introduction to observability with OpenTelemetry
What each signal is for, how they fit together through trace ids, and how to instrument a Python service with OpenTelemetry without drowning in data.
- Tech articles
Caching for backend engineers: cache-aside, TTLs and the stampede problem
The caching patterns you will actually use, how to pick TTLs, how to invalidate safely, and how to stop a cache miss from turning into a thundering herd on your database.
- Tech articles
Composite indexes in SQL: why column order decides everything
How a multi-column B-tree index is actually used, the leftmost-prefix rule, and a simple way to choose column order for filters, ranges and sorting, with EXPLAIN examples.
Share a tech article or a tool you built. Every post is checked and reviewed before it goes live.