I spun up eight threads for a token-parsing job a while back, expecting something close to an 8x speedup on an 8-core box. Watched htop instead, and every core sat idle except one, pegged at 100%, while the other seven did essentially nothing. Total wall-clock time came out worse than if I’d just written it single-threaded and walked away.
Python doesn’t get slow because of the language’s syntax. It gets slow because CPython — the interpreter basically everyone means when they say “Python” — enforces a single lock around bytecode execution, and that lock doesn’t care how many threads are standing in line waiting for a turn. The Global Interpreter Lock exists because CPython’s reference counting isn’t thread-safe on its own; without a single lock serializing access, two threads incrementing the same object’s refcount at the same instant is a real, silent data-corruption bug waiting to happen. The GIL trades away real thread-level parallelism for memory safety simple enough to reason about.
Eight threads doing CPU-bound work in pure Python never actually run at the same time. They take turns, and the switching itself costs something — which is how you end up with worse wall-clock time than a single thread would’ve given you, not just the same time.
Multiprocessing vs. Multithreading: The OS View #
Multiprocessing sidesteps the GIL by not sharing anything to be locked over in the first place. Each worker is a genuinely separate OS process, with its own CPython interpreter, its own heap, its own GIL that nothing else touches. Spin up four processes and you get four fully independent interpreters, each free to run bytecode at the same instant as the others, on separate cores, with nobody waiting on anybody else’s lock.
That independence isn’t free. Creating a process this way costs more than spinning up a thread — fork on Unix is cheap because it copies the parent’s memory via copy-on-write rather than duplicating it outright, while spawn — the default on Windows and macOS since Python 3.8 — starts a genuinely fresh interpreter from scratch, safer and more portable but noticeably slower to set up. And since the processes share nothing, getting data between them means serializing it, usually through pickle , shipping it across an IPC pipe, and deserializing it on the other end. For a few large NumPy arrays passed once, that’s fine. For a tight loop passing small objects back and forth constantly, the serialization overhead can eat the entire benefit of going parallel in the first place.
Crossing the Native Boundary: C Extensions and the C-API #
There’s a third path that neither pure threading nor multiprocessing takes, and it’s the one that makes libraries like NumPy fast without needing a process pool: releasing the GIL from inside a C extension.
CPython’s own C-API has macros built for exactly this — Py_BEGIN_ALLOW_THREADS and Py_END_ALLOW_THREADS , wrapped around a block of C code that doesn’t touch Python objects. Inside that block, the extension has explicitly told the interpreter nothing in there needs the lock, and the GIL gets released for the duration. Other Python threads can run real bytecode during that window instead of queuing behind a lock some C function was holding for no reason. That’s the actual mechanism behind why `numpy.dot()` or a well-written Rust extension can use multiple cores from ordinary Python threads, while a pure-Python loop never can.
The other piece of this is memory. C extensions that allocate outside the normal CPython heap — raw pointers, or a buffer exposed through PyMemoryView — let multiple threads read and write the same data directly, with no pickling and no IPC pipe involved. Threading was never the problem. Holding a lock around code that never needed to hold it was the problem, and native extensions are the fix that goes around that instead of avoiding threading entirely.
Real-World Architecture Matrix #
| Execution Strategy | Primary Bottleneck Solved | Memory Footprint | IPC / Overhead | Ideal Use Case |
|---|---|---|---|---|
| Standard Threading (threading) | I/O waiting (network/disk) | Minimal — shared memory, one interpreter | Extremely low | Async API calls, web scraping, socket listening |
| Multiprocessing (multiprocessing) | CPU-bound work, multi-core | High — duplicate interpreters per worker | High — pickle + IPC on every transfer | Heavy batch data processing, image transforms |
| Native C/Rust Extensions | GIL contention itself | Low — shared process memory, single interpreter | Minimal — direct memory access, no serialization | Numerical computing, cryptography, compression — anywhere a well-written native library already exists |
System Verification: Prove It, Don’t Assume It #
None of the above matters if the actual setup in front of you isn’t doing what you think it’s doing. A library claiming to release the GIL, a thread count that looks parallel on paper — neither is real until the cores have actually done something about it.
Running on X logical cores, X workers per test.
--- Single-threaded (baseline) ---
Wall time: X.XXs
Per-core usage sample: [...]
Cores over 50% utilization: X / X
--- Threaded (pure Python -- GIL bound) ---
Wall time: X.XXs
Per-core usage sample: [...]
Cores over 50% utilization: X / X
--- Multiprocessing (separate interpreters) ---
Wall time: X.XXs
Per-core usage sample: [...]
Cores over 50% utilization: X / XRun it and watch the middle section specifically. The threaded run should land close to the same single-core utilization as the baseline, because pure Python holding the GIL doesn’t buy real parallelism no matter how many threads get spun up. The multiprocessing run is where the active core count should actually climb, because each process carries its own GIL and nobody’s fighting over the same one anymore.
Threading was never broken. It does exactly what it’s for — waiting on I/O without blocking everything else — and falls apart the second the work turns CPU-bound instead of network-bound, because waiting on a socket and holding a lock around a tight math loop aren’t the same problem wearing different clothes.
Pick multiprocessing when the data’s large and infrequent enough to eat the serialization cost. Pick anative extension when someone’s already built one for exactly your problem. Whichever one gets picked, check the cores before trusting the thread count.
