Community/Tech articles

Caching for backend engineers: cache-aside, TTLs and the stampede problem

The caching patterns you will actually use, how to pick TTLs, how to invalidate safely, and how to stop a cache miss from turning into a thundering herd on your database.

A cache is the cheapest performance win in most backends and one of the easiest ways to serve wrong data. The difference comes down to a handful of decisions: which pattern you use, how long entries live, how you invalidate them, and what happens when a popular key expires.

The patterns

Cache-aside (lazy loading) is the default for most services:

def get_product(product_id):
    key = f"product:{product_id}"
    cached = redis.get(key)
    if cached is not None:
        return json.loads(cached)
    product = db.fetch_product(product_id)          # miss: go to the source
    redis.set(key, json.dumps(product), ex=300)     # store for 5 minutes
    return product

The application owns the logic, the cache only holds what was actually requested, and if the cache is down you still work (more slowly).

Write-through updates the cache at the same time as the database. Reads are always warm, but you pay on every write and cache data nobody may read.

Write-behind writes to the cache and flushes to the database later. It's fast but risky, because data can be lost if the cache fails before flushing. Use it only when you understand that trade-off.

For most CRUD services: cache-aside for reads, and delete the key on write.

Invalidate by deleting, not updating

When data changes, delete the cache entry rather than writing the new value into it:

def update_product(product_id, fields):
    db.update_product(product_id, fields)
    redis.delete(f"product:{product_id}")

Writing the new value into the cache seems more efficient, but two concurrent updates can land in the cache in the wrong order and leave a stale value there indefinitely. Deleting means the next read loads the truth from the database.

Even with deletes there's a small race: a reader loads old data from the database just before your update, then writes it to the cache just after your delete. A TTL is your safety net. It caps how long any inconsistency can last.

Choosing TTLs

Ask: how stale can this be before someone notices or something breaks?

Data Typical TTL
Static reference data (countries, plans) hours to a day
Product pages, public profiles 1 to 10 minutes
Counters, "views", leaderboards 10 to 60 seconds
Permissions, prices at checkout don't cache, or a few seconds with explicit invalidation

Add jitter. If you warm 10,000 keys at deploy time with the same TTL, they all expire together. Randomise by around 10%:

ttl = 300 + random.randint(0, 30)

The stampede (thundering herd)

A hot key expires. A thousand requests arrive in the same second, all miss, and all hit the database with the same expensive query. The cache didn't protect you at the moment you needed it most.

Three defences, from simplest to strongest:

  1. Request coalescing with a lock. The first request to miss takes a short lock and rebuilds the value; the others wait briefly and re-read the cache.
def get_with_lock(key, load, ttl=300):
    value = redis.get(key)
    if value is not None:
        return json.loads(value)
    if redis.set(f"lock:{key}", "1", nx=True, ex=10):   # only one rebuilder
        try:
            value = load()
            redis.set(key, json.dumps(value), ex=ttl)
            return value
        finally:
            redis.delete(f"lock:{key}")
    time.sleep(0.05)                                     # others wait, then retry
    return get_with_lock(key, load, ttl)

(In production, cap the retries so a slow rebuild can't pile up waiting requests.)

  1. Serve stale while revalidating. Store the value with a "soft" expiry inside it and a longer hard TTL. After the soft expiry, return the stale value immediately and refresh it in the background.

  2. Early probabilistic refresh. Each read has a small chance of refreshing the key shortly before it expires, which spreads rebuilds out over time.

Cache the negative result too

If a lookup returns "not found", cache that briefly as well (say 30 to 60 seconds). Otherwise a burst of requests for a missing id goes straight to the database every time, which also makes a nice denial-of-service vector.

What not to cache

  • Anything where a stale value causes harm: balances, stock at checkout, authorization decisions.
  • Per-user data in a shared cache under a key that doesn't include the user id. That's how one user sees another user's page.
  • Huge objects you only ever read part of. Cache the part.

Measure it

Track the hit ratio, latency for hits and misses, evictions, and memory. A cache with a 20% hit rate is mostly overhead; one with 99% hits and frequent evictions needs more memory or shorter keys.

Start with cache-aside, delete on write, sensible TTLs with jitter, and lock-based rebuilds for your hottest keys. That covers the vast majority of real systems.

Written by

RecallRun Editors

Practical guides and independent tool overviews from the RecallRun team. Every post is written to be tested on your own machine.

Website

Written by RecallRun Editors for the RecallRun community. Community posts are checked for safety and reviewed by our editors before publishing, but the views and claims are the author's own. Links are the author's; open them with care. Report this post.

More from the community

Write for RecallRun

Share a tech article or a tool you built. Every post is checked and reviewed before it goes live.

Start writing