Vector embeddings explained for developers: similarity, normalisation and the mistakes that hurt search
What an embedding really is, how cosine similarity and dot product relate, why you should normalise, and the practical mistakes that quietly make semantic search worse.
An embedding is a list of numbers, typically a few hundred to a few thousand of them, that represents the meaning of a piece of text. Texts with similar meaning end up close together in that space. That one property powers semantic search, recommendations, clustering, deduplication and the retrieval step of RAG.
You don't need the maths of how the model is trained to use embeddings well. You do need to understand distance, and the handful of mistakes that make results quietly worse.
From text to vector
from openai import OpenAI
client = OpenAI()
resp = client.embeddings.create(
model="text-embedding-3-small",
input=["How do I reset my password?", "Steps to recover account access"],
)
a, b = (d.embedding for d in resp.data)
print(len(a)) # dimension of the vector, fixed per model
Any embedding model works the same way: text in, fixed-length vector out. Open-source models (for example through sentence-transformers or a local runtime) follow the same pattern.
Measuring "closeness"
Three measures come up constantly:
- Cosine similarity: the cosine of the angle between two vectors, from -1 to 1. It ignores vector length and compares direction only.
- Dot product: the sum of element-wise products. It depends on both direction and length.
- Euclidean (L2) distance: straight-line distance.
The key fact: once vectors are normalised to length 1, all three agree on the ranking. For unit vectors, dot product equals cosine similarity, and L2 distance is a simple function of it.
import numpy as np
def normalise(v):
v = np.asarray(v, dtype=np.float32)
return v / np.linalg.norm(v)
a, b = normalise(a), normalise(b)
cosine = float(a @ b) # same as dot product now
So normalise once when you store vectors, and then use the cheapest operation (the dot product) everywhere. Many models already return normalised vectors; check the documentation, or just normalise anyway, since it's cheap.
Searching many vectors
For up to a few hundred thousand vectors, exact search is often fast enough:
docs = np.vstack([normalise(v) for v in doc_vectors]) # shape (n, d)
query = normalise(query_vector)
scores = docs @ query # one matrix-vector product
top = np.argsort(-scores)[:10]
Beyond that, approximate nearest neighbour (ANN) indexes such as HNSW trade a little recall for much faster queries. Databases like PostgreSQL with pgvector, or dedicated vector databases, provide them. Measure recall against exact search on your own data before trusting an index's settings.
Mistakes that hurt search quality
1. Mixing models. Vectors from different models, or even different versions of one model, live in different spaces. Comparing them gives nonsense. Store the model name next to every vector and re-embed everything when you switch.
2. Embedding chunks that are too big. One vector has to summarise the whole chunk. A 3,000-word chunk covering five topics becomes an average that matches none of them well. Split documents along their natural structure (headings, sections) into a few hundred words each.
3. Losing context when chunking. A chunk that says "Set it to 30 seconds" is useless without knowing what "it" is. Prepend the document title and section heading to each chunk before embedding.
4. Treating similarity scores as probabilities. A cosine of 0.82 doesn't mean "82% relevant", and good thresholds vary between models and between corpora. If you need a cut-off, pick it from a labelled sample of your own data.
5. Ignoring exact matches. Embeddings are weak at exact tokens such as error codes, function names and product SKUs. Hybrid search, combining keyword (BM25) and vector results, usually beats either alone.
6. Not evaluating. Write down 30 to 50 real queries with the documents that should come back, and measure hit rate at 1 and at 5 every time you change the model, chunking or index. Without that, tuning is guesswork.
Beyond search
The same vectors are useful for:
- Deduplication: flag pairs above a high similarity threshold.
- Clustering: group support tickets or feedback by topic with k-means or HDBSCAN.
- Classification: a simple logistic regression on embeddings is a strong, cheap baseline.
- Recommendations: "more like this" using the item's own vector as the query.
The essentials
Use one model per index, normalise and use the dot product, chunk by structure with context attached, combine with keyword search for exact terms, and keep a small evaluation set so every change is measured rather than guessed.
Written by RecallRun Editors for the RecallRun community. Community posts are checked for safety and reviewed by our editors before publishing, but the views and claims are the author's own. Links are the author's; open them with care. Report this post.
More from the community
- Tools
pgvector: vector search inside the Postgres you already run
pgvector adds a vector type, distance operators and approximate indexes to PostgreSQL. Here is the setup, the queries, the index choices and when a separate vector database is still worth it.
- Tech articles
Prompt injection: practical defences for LLM apps that read untrusted text
Any LLM feature that reads emails, web pages or user uploads can be steered by text hidden inside them. Here is a layered set of defences that work in production, and why no single one is enough.
- Tools
Ollama: run open LLMs locally for development, privacy and offline work
Ollama makes running open models on your own machine a one-command job and exposes them through a local API. Here is the workflow, how to call it from code, and what to expect on real hardware.
Share a tech article or a tool you built. Every post is checked and reviewed before it goes live.