Community/Tech articles

Vector embeddings explained for developers: similarity, normalisation and the mistakes that hurt search

What an embedding really is, how cosine similarity and dot product relate, why you should normalise, and the practical mistakes that quietly make semantic search worse.

An embedding is a list of numbers, typically a few hundred to a few thousand of them, that represents the meaning of a piece of text. Texts with similar meaning end up close together in that space. That one property powers semantic search, recommendations, clustering, deduplication and the retrieval step of RAG.

You don't need the maths of how the model is trained to use embeddings well. You do need to understand distance, and the handful of mistakes that make results quietly worse.

From text to vector

from openai import OpenAI

client = OpenAI()
resp = client.embeddings.create(
    model="text-embedding-3-small",
    input=["How do I reset my password?", "Steps to recover account access"],
)
a, b = (d.embedding for d in resp.data)
print(len(a))   # dimension of the vector, fixed per model

Any embedding model works the same way: text in, fixed-length vector out. Open-source models (for example through sentence-transformers or a local runtime) follow the same pattern.

Measuring "closeness"

Three measures come up constantly:

  • Cosine similarity: the cosine of the angle between two vectors, from -1 to 1. It ignores vector length and compares direction only.
  • Dot product: the sum of element-wise products. It depends on both direction and length.
  • Euclidean (L2) distance: straight-line distance.

The key fact: once vectors are normalised to length 1, all three agree on the ranking. For unit vectors, dot product equals cosine similarity, and L2 distance is a simple function of it.

import numpy as np

def normalise(v):
    v = np.asarray(v, dtype=np.float32)
    return v / np.linalg.norm(v)

a, b = normalise(a), normalise(b)
cosine = float(a @ b)          # same as dot product now

So normalise once when you store vectors, and then use the cheapest operation (the dot product) everywhere. Many models already return normalised vectors; check the documentation, or just normalise anyway, since it's cheap.

Searching many vectors

For up to a few hundred thousand vectors, exact search is often fast enough:

docs = np.vstack([normalise(v) for v in doc_vectors])   # shape (n, d)
query = normalise(query_vector)
scores = docs @ query                                    # one matrix-vector product
top = np.argsort(-scores)[:10]

Beyond that, approximate nearest neighbour (ANN) indexes such as HNSW trade a little recall for much faster queries. Databases like PostgreSQL with pgvector, or dedicated vector databases, provide them. Measure recall against exact search on your own data before trusting an index's settings.

Mistakes that hurt search quality

1. Mixing models. Vectors from different models, or even different versions of one model, live in different spaces. Comparing them gives nonsense. Store the model name next to every vector and re-embed everything when you switch.

2. Embedding chunks that are too big. One vector has to summarise the whole chunk. A 3,000-word chunk covering five topics becomes an average that matches none of them well. Split documents along their natural structure (headings, sections) into a few hundred words each.

3. Losing context when chunking. A chunk that says "Set it to 30 seconds" is useless without knowing what "it" is. Prepend the document title and section heading to each chunk before embedding.

4. Treating similarity scores as probabilities. A cosine of 0.82 doesn't mean "82% relevant", and good thresholds vary between models and between corpora. If you need a cut-off, pick it from a labelled sample of your own data.

5. Ignoring exact matches. Embeddings are weak at exact tokens such as error codes, function names and product SKUs. Hybrid search, combining keyword (BM25) and vector results, usually beats either alone.

6. Not evaluating. Write down 30 to 50 real queries with the documents that should come back, and measure hit rate at 1 and at 5 every time you change the model, chunking or index. Without that, tuning is guesswork.

The same vectors are useful for:

  • Deduplication: flag pairs above a high similarity threshold.
  • Clustering: group support tickets or feedback by topic with k-means or HDBSCAN.
  • Classification: a simple logistic regression on embeddings is a strong, cheap baseline.
  • Recommendations: "more like this" using the item's own vector as the query.

The essentials

Use one model per index, normalise and use the dot product, chunk by structure with context attached, combine with keyword search for exact terms, and keep a small evaluation set so every change is measured rather than guessed.

Written by

RecallRun Editors

Practical guides and independent tool overviews from the RecallRun team. Every post is written to be tested on your own machine.

Website

Written by RecallRun Editors for the RecallRun community. Community posts are checked for safety and reviewed by our editors before publishing, but the views and claims are the author's own. Links are the author's; open them with care. Report this post.

More from the community

Write for RecallRun

Share a tech article or a tool you built. Every post is checked and reviewed before it goes live.

Start writing