
01 · What Are Embeddings?
Dense vector representations turn discrete objects into comparable coordinates in a learned space. Start here for the core idea: useful similarity becomes geometry.
Foundations
A visual guide to the representation layer behind semantic search, recommendations, vector databases, and RAG: plus the production choices that determine quality, latency, cost, and governance.

Without embeddings
Keyword systems can miss conceptually related content when the vocabulary changes. Large sparse feature spaces also make many learning and retrieval tasks cumbersome.
With embeddings
Queries and content can be mapped into a learned vector space, making semantic similarity, clustering, recommendation, and retrieval practical.
The embedding model is only one part of retrieval quality. Chunking, metadata, query formulation, vector index, hybrid retrieval, reranking, and evaluation can matter just as much.
Complete visual guide
Use the slides for the mental model; use the production notes below for current engineering guidance. Click any slide to enlarge.
Key concepts
A numeric representation learned so that similarity or task relationships become accessible through vector operations.
Encodes queries and documents independently so document vectors can be precomputed and searched efficiently.
Applies a more expensive relevance model to a smaller candidate set after first-stage retrieval.
Combines dense semantic retrieval with sparse lexical signals to retain both meaning and exact-term precision.
Approximate-nearest-neighbor structure that trades some exactness for practical search speed at scale.
A change in vector space or corpus distribution that can invalidate assumptions about comparability and retrieval quality.
From tutorial to production
Clean, permission, deduplicate, enrich metadata, and chunk the source corpus.
Select model, instruction format, dimension, normalization, batching, and version.
Select vector store, index, metric, filters, hybrid strategy, top-k, and reranking.
Measure relevance, latency, cost, failure slices, and downstream answer quality.
MTEB is now a broad, multilingual embedding benchmark and leaderboard, useful for shortlisting models. OpenAI's current v3 embedding models include text-embedding-3-small and text-embedding-3-large; text-embedding-ada-002 should be treated as an older model rather than the default current reference.
Build the document, embedding, index, retrieval, reranking, and RAG pipeline.
Govern source access, model/version lineage, evaluation evidence, and rollout approvals.
Make chunk size, model, dimensions, top-k, filters, thresholds, index settings, and reranking measurable and adjustable.
FAQ
Embeddings are numeric vector representations of content. A useful embedding space places items that are related for a task near each other, enabling similarity search, clustering, recommendation, classification, and retrieval.
Classic word embeddings assign vectors to words, while sentence or passage embeddings represent a larger unit of meaning and are usually the more practical choice for modern semantic search and RAG retrieval.
Cosine similarity compares vector direction rather than raw magnitude. It is common for normalized text embeddings, but the correct similarity metric should follow the embedding model and vector index documentation.
Documents are chunked and embedded into an index. The user query is embedded, similar chunks are retrieved, and those chunks are provided to the language model as evidence for generation.
Use a leaderboard such as MTEB to create a shortlist, then evaluate the candidates on your own corpus, languages, query patterns, metadata filters, latency target, cost, and downstream answer quality.
Fine-tune after establishing a strong general baseline and a reliable evaluation set. Domain-specific query-document pairs and hard negatives are especially useful when vocabulary and relevance patterns differ from general web text.
