Inside CPython: bytecode & memory
"Python" is a language. CPython is the program, written in C, that runs it when you type
python. It compiles your source to bytecode, executes that bytecode in a
loop, and manages memory with reference counting plus a cycle collector. Knowing this
explains performance advice, __pycache__ folders, and why del does not always
free memory.
A score and an orchestra
Your .py file is a composer's draft. The compiler turns it into a score of simple
instructions (bytecode). The interpreter is a tireless player who reads the score one instruction at a
time: "fetch this, add those, call that". It never sees your original draft again. Faster Pythons
mostly make the player smarter: PyPy, the 3.11+ specialising interpreter, and the experimental JIT.
1. From source to execution
2. Reading bytecode with dis
CPython is a stack machine: instructions push values onto a stack and pop them off. The
dis module shows the instructions for any function.
import dis
def total_price(price, qty):
return price * qty + 5
dis.dis(total_price)
Read it top to bottom. One fused instruction pushes both price and qty onto the
stack (3.14 combines common pairs of loads). BINARY_OP (*) pops two values and pushes their
product. LOAD_SMALL_INT pushes 5, BINARY_OP (+) adds, and
RETURN_VALUE returns the top of the stack. RESUME is a bookkeeping instruction at
the start of every function.
Instruction names and layout change in almost every Python release. The output above is from
CPython 3.14, and 3.12 or 3.13 show different instructions for the same function. Use
dis to understand and compare code. Never write code that depends on the exact bytecode.
def total_price(price, qty):
discount = 5
return price * qty + discount
code = total_price.__code__
print("locals: ", code.co_varnames)
print("constants:", code.co_consts)
print("arg count:", code.co_argcount)
print("bytes of bytecode:", len(code.co_code))
3. Why Python got faster: the specialising interpreter
Since 3.11, CPython specialises instructions at run time. After a BINARY_OP has seen two
ints a few times, it is rewritten in place to a faster int-only version, with a guard that falls back if
the types change. That is a large part of the 10–60% speedups in 3.11. Two newer options are the
free-threaded build (no GIL, officially supported since 3.14; lesson 16) and an experimental
JIT compiler that can be enabled in some builds. Both are still optional.
import dis
def add(a, b):
return a + b
for _ in range(1000): # warm it up with ints
add(1, 2)
dis.dis(add, adaptive=True) # shows the specialised instructions now in use
4. Memory: reference counting
Every object carries a count of how many references point to it. Binding a name, putting the object in a list, or passing it to a function increments the count. When the count drops to zero, CPython frees the object immediately. That is why files closed by a function going out of scope usually close at a predictable moment in CPython. PyPy and other implementations do not promise this.
import sys
class Tracked:
def __del__(self):
print(" freed!")
data = Tracked()
print("refcount:", sys.getrefcount(data)) # +1 for getrefcount's own argument
alias = data
print("refcount:", sys.getrefcount(data))
print("del data")
del data # removes a NAME, not the object; alias still refers to it
print("del alias")
del alias # last reference gone, so it is freed right now
print("done")
del deletes names, not objects
del x unbinds the name x and decrements the object's count. The object is freed
only when no references remain. If a list, cache or closure still holds one, del frees nothing.
5. The cycle collector
Reference counting cannot free cycles: two objects that refer to each other keep each other's
count above zero forever. CPython's gc module periodically finds groups of container
objects that are reachable only from each other and frees them.
import gc
class Node:
def __init__(self, name):
self.name, self.other = name, None
def __del__(self):
print(" freed", self.name)
gc.disable() # so we control when collection happens
a, b = Node("a"), Node("b")
a.other, b.other = b, a # a cycle
del a, b
print("names deleted; nothing freed yet, the cycle keeps both alive")
print("collected:", gc.collect(), "unreachable objects")
The collector is generational: new objects are checked often, and survivors are checked less
often. For almost every program the defaults are right. gc.freeze() and tuning thresholds
matter mainly for very large, long-lived services.
6. Caching and interning
import sys
a = int("256"); b = int("256")
print(a is b) # small ints (-5..256) are cached singletons in CPython
a = int("257"); b = int("257")
print(a is b) # outside the cache: two separate objects
s1 = "".join(["data", "_", "lake"])
s2 = "".join(["data", "_", "lake"])
print(s1 == s2, s1 is s2)
print(sys.intern(s1) is sys.intern(s2)) # interning maps equal strings to one object
int("…") and "".join are used here to build values at run time. Literals written
in the same file can be merged by the compiler, which would make the is results misleading.
That is exactly why is must not be used to compare numbers or strings.
7. How big are objects?
import sys
for value in [0, 1, 2**30, 2**100, 1.0, "", "a", "abc", (), [], {}, set()]:
print(f"{value!r:>32} {sys.getsizeof(value):>4} bytes")
nums = list(range(1000))
print("list of 1000 ints:", sys.getsizeof(nums), "bytes for the list itself,",
sum(map(sys.getsizeof, nums)), "for the int objects it points to")
sys.getsizeof reports only the object itself, not what it points to. A list stores
pointers (8 bytes each on 64-bit) to separate int objects. That is why a NumPy array of a million
numbers uses a fraction of the memory of a Python list of the same numbers.
Recap
- Source → tokens → AST → bytecode → eval loop. Imported modules are cached as
.pyc. - CPython is a stack machine;
disshows the instructions, which change between versions. - 3.11+ specialises instructions at run time. Free-threading and a JIT are optional newer builds.
- Reference counting frees objects immediately;
delremoves a name, not an object. - The cycle collector cleans up objects that only reference each other.
- Small ints and interned strings are shared, an implementation detail you must not rely on.
Checkpoint
del big_list but memory does not drop. What is the most likely reason?
del only unbinds one name. The object is freed when its last reference disappears. Look for aliases, caches such as lru_cache, or objects held by tracebacks and frames.a.partner = b; b.partner = a after del a, b?
gc detects groups of objects that are unreachable from outside and frees them.