Blog
Experiments on RAG, agents, MCP and evals. Every number on this site comes from code that was actually run.
6 experiments publishedScript + raw output for every measurementNew posts Mon · Wed · Fri
Do you need a vector database? Brute-force search benchmarked from 10k to 1M vectors
Exact numpy search over 100k embeddings takes 3.5 ms on 2 vCPUs. Measured latency, memory and HNSW recall to show when an index actually pays off.
Hybrid search in Python: BM25 + vectors with Reciprocal Rank Fusion, alpha swept 0 to 1
Tested BM25+RRF on 2,713 FastAPI doc chunks: BM25 beat vectors by 49 MRR points on identifiers; vectors beat BM25 by 24 on paraphrases. alpha=0.8 balanced both.
Jev vs LLMs: when a decision model beats a chat model (and when it doesn't)
Analysis: Jev claims 70-500ms typed decisions (vendor-reported). My runnable stand-in, TF-IDF + logistic regression, routed text at 0.6ms and $0, 60% accuracy.
- Newsletter
Get each new experiment by email
One short email per post with the numbers, plus the free 25-point RAG production checklist.
RAG chunking, measured: heading-aware chunks doubled hit@1 on the FastAPI docs
I tested 5 chunking strategies on 151 real docs and 773 queries. Fixed-size chunks crossed section boundaries 59% of the time. Heading-aware chunks never did.
The RAG production checklist: 25 checks before real users see it
25 checks across ingestion, retrieval, generation, evaluation and operations that catch the RAG failures demos hide. Free printable PDF included.
Stop fake citations in RAG: validate them server-side in 25 lines of Python
LLMs will cite passage [7] when you gave them 5. A tested validator that strips invented citations and returns a grounded flag your UI can trust.