Module 1 · How Python really works

Inside CPython: bytecode & memory

Advanced 20 min read What actually runs your code

"Python" is a language. CPython is the program, written in C, that runs it when you type python. It compiles your source to bytecode, executes that bytecode in a loop, and manages memory with reference counting plus a cycle collector. Knowing this explains performance advice, __pycache__ folders, and why del does not always free memory.

🎼

A score and an orchestra

Your .py file is a composer's draft. The compiler turns it into a score of simple instructions (bytecode). The interpreter is a tireless player who reads the score one instruction at a time: "fetch this, add those, call that". It never sees your original draft again. Faster Pythons mostly make the player smarter: PyPy, the 3.11+ specialising interpreter, and the experimental JIT.

1. From source to execution

Source app.py (text) Tokens NAME, OP, NUMBER AST ast.parse() Code object bytecode + constants Eval loop ceval.c executes __pycache__/app.cpython-314.pyc Imported modules are compiled once and cached as .pyc. The script you run directly is recompiled every time.

2. Reading bytecode with dis

CPython is a stack machine: instructions push values onto a stack and pop them off. The dis module shows the instructions for any function.

import dis

def total_price(price, qty):
    return price * qty + 5

dis.dis(total_price)
3 RESUME 0 4 LOAD_FAST_BORROW_LOAD_FAST_BORROW 1 (price, qty) BINARY_OP 5 (*) LOAD_SMALL_INT 5 BINARY_OP 0 (+) RETURN_VALUE

Read it top to bottom. One fused instruction pushes both price and qty onto the stack (3.14 combines common pairs of loads). BINARY_OP (*) pops two values and pushes their product. LOAD_SMALL_INT pushes 5, BINARY_OP (+) adds, and RETURN_VALUE returns the top of the stack. RESUME is a bookkeeping instruction at the start of every function.

Bytecode is not a stable API

Instruction names and layout change in almost every Python release. The output above is from CPython 3.14, and 3.12 or 3.13 show different instructions for the same function. Use dis to understand and compare code. Never write code that depends on the exact bytecode.

def total_price(price, qty):
    discount = 5
    return price * qty + discount

code = total_price.__code__
print("locals:   ", code.co_varnames)
print("constants:", code.co_consts)
print("arg count:", code.co_argcount)
print("bytes of bytecode:", len(code.co_code))
locals: ('price', 'qty', 'discount') constants: (5,) arg count: 2 bytes of bytecode: 36

3. Why Python got faster: the specialising interpreter

Since 3.11, CPython specialises instructions at run time. After a BINARY_OP has seen two ints a few times, it is rewritten in place to a faster int-only version, with a guard that falls back if the types change. That is a large part of the 10–60% speedups in 3.11. Two newer options are the free-threaded build (no GIL, officially supported since 3.14; lesson 16) and an experimental JIT compiler that can be enabled in some builds. Both are still optional.

import dis

def add(a, b):
    return a + b

for _ in range(1000):      # warm it up with ints
    add(1, 2)

dis.dis(add, adaptive=True)  # shows the specialised instructions now in use
3 RESUME_CHECK 0 4 LOAD_FAST_BORROW_LOAD_FAST_BORROW 1 (a, b) BINARY_OP_ADD_INT 0 (+) RETURN_VALUE

4. Memory: reference counting

Every object carries a count of how many references point to it. Binding a name, putting the object in a list, or passing it to a function increments the count. When the count drops to zero, CPython frees the object immediately. That is why files closed by a function going out of scope usually close at a predictable moment in CPython. PyPy and other implementations do not promise this.

import sys

class Tracked:
    def __del__(self):
        print("  freed!")

data = Tracked()
print("refcount:", sys.getrefcount(data))   # +1 for getrefcount's own argument
alias = data
print("refcount:", sys.getrefcount(data))

print("del data")
del data            # removes a NAME, not the object; alias still refers to it
print("del alias")
del alias           # last reference gone, so it is freed right now
print("done")
refcount: 2 refcount: 3 del data del alias freed! done
del deletes names, not objects

del x unbinds the name x and decrements the object's count. The object is freed only when no references remain. If a list, cache or closure still holds one, del frees nothing.

5. The cycle collector

Reference counting cannot free cycles: two objects that refer to each other keep each other's count above zero forever. CPython's gc module periodically finds groups of container objects that are reachable only from each other and frees them.

import gc

class Node:
    def __init__(self, name):
        self.name, self.other = name, None
    def __del__(self):
        print("  freed", self.name)

gc.disable()                  # so we control when collection happens
a, b = Node("a"), Node("b")
a.other, b.other = b, a       # a cycle
del a, b
print("names deleted; nothing freed yet, the cycle keeps both alive")
print("collected:", gc.collect(), "unreachable objects")
names deleted; nothing freed yet, the cycle keeps both alive freed a freed b collected: 2 unreachable objects

The collector is generational: new objects are checked often, and survivors are checked less often. For almost every program the defaults are right. gc.freeze() and tuning thresholds matter mainly for very large, long-lived services.

6. Caching and interning

import sys

a = int("256"); b = int("256")
print(a is b)            # small ints (-5..256) are cached singletons in CPython

a = int("257"); b = int("257")
print(a is b)            # outside the cache: two separate objects

s1 = "".join(["data", "_", "lake"])
s2 = "".join(["data", "_", "lake"])
print(s1 == s2, s1 is s2)
print(sys.intern(s1) is sys.intern(s2))   # interning maps equal strings to one object
True False True False True

int("…") and "".join are used here to build values at run time. Literals written in the same file can be merged by the compiler, which would make the is results misleading. That is exactly why is must not be used to compare numbers or strings.

7. How big are objects?

import sys

for value in [0, 1, 2**30, 2**100, 1.0, "", "a", "abc", (), [], {}, set()]:
    print(f"{value!r:>32}  {sys.getsizeof(value):>4} bytes")

nums = list(range(1000))
print("list of 1000 ints:", sys.getsizeof(nums), "bytes for the list itself,",
      sum(map(sys.getsizeof, nums)), "for the int objects it points to")
0 28 bytes 1 28 bytes 1073741824 32 bytes 1267650600228229401496703205376 40 bytes 1.0 24 bytes '' 41 bytes 'a' 42 bytes 'abc' 44 bytes () 48 bytes [] 56 bytes {} 64 bytes set() 216 bytes list of 1000 ints: 8056 bytes for the list itself, 28000 for the int objects it points to

sys.getsizeof reports only the object itself, not what it points to. A list stores pointers (8 bytes each on 64-bit) to separate int objects. That is why a NumPy array of a million numbers uses a fraction of the memory of a Python list of the same numbers.

Recap

  • Source → tokens → AST → bytecode → eval loop. Imported modules are cached as .pyc.
  • CPython is a stack machine; dis shows the instructions, which change between versions.
  • 3.11+ specialises instructions at run time. Free-threading and a JIT are optional newer builds.
  • Reference counting frees objects immediately; del removes a name, not an object.
  • The cycle collector cleans up objects that only reference each other.
  • Small ints and interned strings are shared, an implementation detail you must not rely on.

Checkpoint

1 · You del big_list but memory does not drop. What is the most likely reason?
del only unbinds one name. The object is freed when its last reference disappears. Look for aliases, caches such as lru_cache, or objects held by tracebacks and frames.
2 · Why can't reference counting alone free a.partner = b; b.partner = a after del a, b?
A cycle keeps every member's count at least 1. CPython's generational gc detects groups of objects that are unreachable from outside and frees them.