Blog/vector-search/Do you need a vector database? Brute-force search benchmarked from 10k to 1M vectors

Do you need a vector database? Brute-force search benchmarked from 10k to 1M vectors

Exact numpy search over 100k embeddings takes 3.5 ms on 2 vCPUs. Measured latency, memory and HNSW recall to show when an index actually pays off.

What this post covers
  1. The setup
  2. Exact search latency
  3. What an HNSW index buys you
  4. Where the time actually goes in a RAG request
  5. My rule of thumb
  6. Reproduce it
How this was testedhardware, versions, data

2 vCPU Intel Xeon @ 2.1 GHz, 7 GB RAM, Python 3.11, numpy 2.4.4, faiss-cpu 1.15.1, hnswlib. Synthetic normalized float32 vectors. Median of 50 (exact) or 200 (HNSW) single queries.

Most RAG tutorials start by installing a vector database. For a lot of projects, that is a dependency you don’t need yet. I measured how long plain exact search takes as the corpus grows, and what an HNSW index buys you in return.

Short answer: below about 100,000 chunks, a numpy matrix multiply is fast enough for almost any RAG app. Between 100k and 1M it depends on your traffic. Past that, use an index.

The setup

Every embedding is normalized, so cosine similarity is a dot product. Exact search is one matrix-vector product and a partial sort:

import numpy as np

def top_k(X: np.ndarray, q: np.ndarray, k: int = 10) -> np.ndarray:
    """X: (n, d) normalized float32 matrix. q: (d,) normalized query."""
    scores = X @ q
    idx = np.argpartition(-scores, k)[:k]
    return idx[np.argsort(-scores[idx])]

I timed this against faiss.IndexFlatIP (also exact) for three common embedding sizes: 384 (MiniLM-class models), 768 (BERT-base-class) and 1536 (OpenAI text-embedding-3-small default). The machine is deliberately small: 2 vCPUs, the size of a cheap cloud instance.

Exact search latency

Single-query latency, median of 50 queries:

Dim Vectors RAM for vectors numpy faiss Flat
384 10,000 15 MB 0.35 ms 0.78 ms
384 100,000 154 MB 3.5 ms 8.3 ms
384 500,000 768 MB 31 ms 66 ms
384 1,000,000 1.5 GB 62 ms 122 ms
768 100,000 307 MB 10.6 ms 22.6 ms
768 500,000 1.5 GB 57 ms 120 ms
768 1,000,000 3.1 GB 114 ms 240 ms
1536 10,000 61 MB 1.2 ms 3.1 ms
1536 100,000 614 MB 20 ms 51 ms
1536 500,000 3.1 GB 107 ms 248 ms

Three things stand out:

  1. Latency scales linearly with n × d. It is a memory-bandwidth problem. Doubling either the corpus or the dimension doubles the time.
  2. Memory is the real limit, not speed. 1M vectors at 1536 dims is 6 GB of float32 before any metadata. I skipped that row: it won’t fit comfortably on a 7 GB machine.
  3. Plain numpy beat faiss Flat for single queries here, by about 2×. Faiss is built for batches; one query at a time pays overhead that numpy’s direct matrix-vector call doesn’t. If you batch queries, test again before assuming this holds.

What an HNSW index buys you

Next, I built hnswlib indexes (M=16, ef_construction=200) and measured query time and recall@10: the fraction of the true top 10 the index actually returns.

I used two synthetic datasets, because recall depends heavily on how your data is shaped:

  • Uniform: random Gaussian directions. This is the worst case for any approximate index: no structure to exploit.
  • Clustered: 2,000 tight clusters. Real embeddings sit between these two, usually much closer to clustered, since documents about the same topic land near each other.
Data Dim Vectors Build time ef=32 ef=64 ef=128
clustered 384 100k 23 s 0.10 ms, recall 0.954 0.17 ms, 0.995 0.34 ms, 1.000
clustered 768 100k 48 s 0.22 ms, 0.878 0.36 ms, 0.985 0.68 ms, 1.000
uniform 384 100k 37 s 0.16 ms, 0.129 0.31 ms, 0.206 0.55 ms, 0.281
uniform 768 100k 69 s 0.34 ms, 0.058 0.63 ms, 0.105 1.07 ms, 0.172
uniform 384 500k 279 s 0.19 ms, 0.034 0.35 ms, 0.057 0.67 ms, 0.099

HNSW queries are 10 to 100 times faster than exact search at these sizes. The cost:

  • Build time. 500k vectors took over 4.5 minutes on 2 cores. You pay this again on every full re-index, such as after switching embedding models.
  • Recall is not guaranteed. On structured data, ef=64 gave near-perfect recall. On structureless data, the same settings returned less than a quarter of the true neighbors. Your data will be somewhere in between, so measure recall on your own embeddings before trusting the defaults.

Where the time actually goes in a RAG request

A typical RAG request spends 1 to 5 seconds waiting for the LLM to generate. Next to that:

  • 3.5 ms (100k × 384, exact) is invisible.
  • 60 to 110 ms (1M exact) is noticeable only if you care about p99, or serve many queries per second on the same box.
  • Exact search gives perfect recall for free, which removes one variable when you debug bad answers.

A rough sizing rule: one PDF page is usually 2 to 3 chunks. 100k chunks is roughly 30,000 to 50,000 pages. Many internal “chat with our docs” projects never get there.

My rule of thumb

Your corpus What I’d use
under 100k chunks numpy (or pgvector with no index). Store vectors in a .npy file or Postgres.
100k to 1M chunks exact search is still fine at low QPS. Add HNSW when p95 latency or QPS demands it, and measure recall.
over 1M chunks, or many tenants a proper ANN index (pgvector HNSW, Qdrant, Milvus, OpenSearch), plus filtering and sharding.

Most importantly, keep the vector store behind one interface so you can swap numpy for pgvector later without touching the rest of the pipeline. That one seam is worth more than picking the “right” database on day one.

Reproduce it

The exact-search benchmark:

import time, numpy as np, faiss
faiss.omp_set_num_threads(2)
rng = np.random.default_rng(0)
norm = lambda x: x / np.linalg.norm(x, axis=1, keepdims=True)
for d in (384, 768, 1536):
    for n in (10_000, 100_000, 500_000, 1_000_000):
        if n * d * 4 > 3.2e9: continue          # skip what won't fit in RAM
        X = norm(rng.standard_normal((n, d), dtype=np.float32))
        Q = norm(rng.standard_normal((50, d), dtype=np.float32))
        t = []
        for q in Q:
            s = time.perf_counter()
            sc = X @ q
            idx = np.argpartition(-sc, 10)[:10]
            t.append(time.perf_counter() - s)
        print(d, n, f"{np.median(t)*1e3:.2f} ms")

For HNSW, build with hnswlib.Index(space="ip", dim=d), init_index(max_elements=n, ef_construction=200, M=16), and compare knn_query results against the exact top 10 to get recall.

Caveat: these are synthetic vectors on one small machine. Latency for exact search doesn’t depend on what the vectors mean, so those numbers transfer well. HNSW recall does depend on your data, which is exactly why you should measure it yourself.

Want the production version of this?

A production RAG + MCP starter kit for FastAPI: hybrid search, validated citations, evals in CI, Docker, 32 tests. It runs offline in about 60 seconds.

See ShipRAG  or get the free RAG checklist first.

Written and measured by

Alby Tomy

Python/FastAPI engineer building RAG and MCP systems. Leads an engineering team in Bengaluru, and runs every experiment on this site before writing about it.

X · GitHub · Contact & enquiries