Blog/rag/Hybrid search in Python: BM25 + vectors with Reciprocal Rank Fusion, alpha swept 0 to 1
Hybrid search in Python: BM25 + vectors with Reciprocal Rank Fusion, alpha swept 0 to 1
Tested BM25+RRF on 2,713 FastAPI doc chunks: BM25 beat vectors by 49 MRR points on identifiers; vectors beat BM25 by 24 on paraphrases. alpha=0.8 balanced both.
What this post covers
How this was tested
FastAPI English docs (151 Markdown files, 2,713 heading-aware chunks), sentence-transformers 6.1.0 with all-MiniLM-L6-v2 (384-dim), Python 3.12.14, faiss-cpu 1.15.1. 773 heading queries, 305 exact-identifier queries, 30 hand-written paraphrase queries. Full script and raw output at the end.
Should you fuse BM25 with vector search, and what weight should each side get? I ran both retrievers separately and fused them with Reciprocal Rank Fusion (RRF) across three different query styles, sweeping the fusion weight from pure vector to pure BM25.
Result: no single retriever wins everywhere. Pure BM25 beat pure vectors by 49 points of MRR on exact code identifiers (0.772 vs 0.279). Pure vectors beat pure BM25 by 24 points on hand-written paraphrase questions (0.768 vs 0.566). An RRF blend with 80% weight on BM25 rank was the only setting that stayed within 9 points of the best score on every query type.
The setup
- Corpus: the English FastAPI docs, 151 Markdown files, chunked with the heading-aware-plus-heading-path chunker from my chunking benchmark — the strategy that scored best there. 2,713 chunks.
- BM25: the same from-scratch implementation as the chunking post (k1=1.5, b=0.75), reused unchanged.
- Vectors:
sentence-transformers/all-MiniLM-L6-v2, 384 dimensions, cosine similarity via normalized dot product. It’s a small, CPU-friendly model, documented on the Hugging Face model card, verified current as of 2026-09-27. - Fusion:
score(d) = alpha / (60 + rank_bm25(d)) + (1 - alpha) / (60 + rank_vec(d)). This is standard RRF withk=60, the constant from the original Cormack, Clarke and Buettcher paper, extended with a weightalphainstead of a flat 50/50 sum.alpha=1.0is pure BM25,alpha=0.0is pure vector search. Each retriever contributes its own top 50 candidates before fusion; a document missing from one side gets a rank of 200 for that side’s term. - Three query sets, on purpose:
- Heading queries (773): section headings used verbatim as queries, same harness as the chunking post. These overlap with the indexed heading-path prefix by construction, so they’re an easy benchmark, not a fair semantic test — included for continuity, not as the main result.
- Identifier queries (305): inline-code spans like
`response_model`or`BackgroundTasks`that appear in exactly one chunk in the whole corpus. This tests exact-token recall, the case hybrid search is usually sold on. - Paraphrase queries (30): questions I wrote by hand — “How do I run code after sending the response to the client?” for the background-tasks page, “How can a client upload a file to my API?” for the file-upload page — deliberately avoiding the target page’s own vocabulary. A hit counts any chunk from the named target file. Small sample, so treat the percentages as approximate, not precise.
Results
MRR by alpha (0.0 = pure vector, 1.0 = pure BM25):
| alpha | heading MRR | identifier MRR | paraphrase MRR |
|---|---|---|---|
| 0.0 | 0.491 | 0.279 | 0.768 |
| 0.2 | 0.581 | 0.471 | 0.794 |
| 0.4 | 0.642 | 0.569 | 0.806 |
| 0.5 | 0.679 | 0.696 | 0.751 |
| 0.6 | 0.710 | 0.735 | 0.731 |
| 0.8 | 0.782 | 0.753 | 0.722 |
| 0.9 | 0.813 | 0.768 | 0.678 |
| 1.0 | 0.859 | 0.772 | 0.566 |
hit@1 / hit@5 for the two more representative sets (identifier and paraphrase):
| alpha | identifier hit@1 | identifier hit@5 | paraphrase hit@1 | paraphrase hit@5 |
|---|---|---|---|---|
| 0.0 | 0.207 | 0.397 | 0.667 | 0.933 |
| 0.4 | 0.515 | 0.656 | 0.667 | 0.967 |
| 0.8 | 0.662 | 0.889 | 0.533 | 0.933 |
| 1.0 | 0.692 | 0.889 | 0.367 | 0.867 |
What the numbers mean
BM25 wins big on identifiers, vectors win big on paraphrases — as expected, but by more than I guessed. MiniLM’s MRR on identifier queries (0.279) is worse than a coin flip at hit@1 (0.207). A 384-dim sentence embedding model wasn’t trained to preserve rare exact tokens like BackgroundTasks; it embeds the gist of the sentence around it, and the gist of one API reference page looks a lot like another. BM25’s inverse-document-frequency weighting does the opposite: a rare token dominates the score.
Pure BM25 also won the heading benchmark, and that number is misleading. The heading-aware chunker (correctly, per the last post) prepends each chunk’s own heading path to the indexed text. Querying with that same heading text is close to a literal substring search, which is exactly BM25’s best case. Real users don’t type section titles verbatim. That’s why the paraphrase set — 30 questions I wrote without copying the target page’s wording — is the fairer read on “does semantic search help,” and there, pure vectors beat pure BM25 by 24 MRR points.
No fixed alpha is best for all three. alpha=0.4 is best for paraphrase, alpha=1.0 is best for identifiers and headings. Computing the regret (best-possible MRR minus this alpha’s MRR) at every alpha and taking the worst of the three query types, the minimum worst-case regret is at alpha=0.8: 1.9 points below the best identifier score, 8.4 below the best paraphrase score and 7.7 below the best heading score. Every other alpha loses more than 8.4 points somewhere: alpha=0.5, the “obvious” 50/50 choice, loses 18 points on headings.
Honest caveats
- The paraphrase set is small and I wrote it. 30 questions is enough to see a 24-point gap, not enough to trust the third decimal place. I picked topics I know FastAPI covers narrowly (one page per topic), which makes the “any chunk from that file counts” hit rule generous. Build a bigger, blinder set before trusting this in production.
- The heading benchmark favors whichever chunker embeds the query text into the index, which is BM25’s easiest case, not a neutral test. I kept it for continuity with the chunking post, not because it represents real usage.
- One small local embedding model. all-MiniLM-L6-v2 is 90 MB and made for exactly this kind of semantic-similarity task — larger or API-hosted embedding models could close some of the identifier gap. I didn’t test that; run this script against your own model if you rely on one.
- RRF’s
k=60wasn’t swept. Only the alpha weight was. A smaller k sharpens the influence of a document’s exact rank; that’s a separate experiment. - Single corpus, single domain. FastAPI’s own docs are unusually well-structured. A messier corpus would likely show a smaller BM25 edge on identifiers, since chunk boundaries would be noisier.
The fusion code
K_RRF = 60
def rrf_fuse(bm25_ranked_ids, vec_ranked_ids, alpha, top_k=5):
"""bm25_ranked_ids, vec_ranked_ids: lists of doc ids, best first.
alpha=1.0 is pure BM25, alpha=0.0 is pure vector."""
rank_bm = {doc_id: r for r, doc_id in enumerate(bm25_ranked_ids)}
rank_vec = {doc_id: r for r, doc_id in enumerate(vec_ranked_ids)}
fallback = max(len(bm25_ranked_ids), len(vec_ranked_ids)) * 4
candidates = set(rank_bm) | set(rank_vec)
scored = [
(
alpha / (K_RRF + rank_bm.get(d, fallback) + 1)
+ (1 - alpha) / (K_RRF + rank_vec.get(d, fallback) + 1),
d,
)
for d in candidates
]
scored.sort(reverse=True)
return [d for _, d in scored[:top_k]]
Run BM25 and your vector index independently, each returning their own top-N (50 is enough here), then fuse. No score normalization needed — that’s the point of using ranks instead of raw scores, which differ in scale between BM25 and cosine similarity.
What I’d do in your pipeline
- Don’t default to alpha=0.5. It’s the “obviously fair” split, but on this corpus it loses more than any other tested value on at least one query type. Start from alpha=0.7-0.8 (BM25-leaning) and tune from a real query log.
- If your users type API names, error codes or config keys, weight BM25 heavily. Small embedding models are bad at exact tokens; that’s structural, not a tuning problem.
- If your users type full questions, don’t skip vectors. BM25 alone lost 24 MRR points against a plain paraphrase set here — a bigger drop than I expected going in.
- Measure your own alpha. The best value depends on how your users actually phrase queries, which the golden-set harness pattern can capture from real traffic instead of a guess.
Full benchmark script
bench_hybrid.py, results.txt and fetch_corpus.sh are in experiments/hybrid-search-bm25-rrf-python/ in this site’s repo. It builds the chunk index, runs BM25 and MiniLM retrieval separately, fuses with RRF at 11 alpha values, and prints hit@1/hit@5/MRR for all three query sets plus median per-query latency (BM25: 1.82ms, vector encode+search: 23.99ms, both on 2 vCPUs over 2,713 chunks) — a hybrid query costs roughly what the vector half costs, since BM25 is nearly free by comparison. ShipRAG ships this same BM25+RRF hybrid search wired into a FastAPI starter, if you’d rather not build the fusion layer from scratch.
Want the production version of this?
A production RAG + MCP starter kit for FastAPI: hybrid search, validated citations, evals in CI, Docker, 32 tests. It runs offline in about 60 seconds.
See ShipRAG or get the free RAG checklist first.