Embeddings & RAG · Chunking and Tokenization

Chunking vs Tokenization: How to Prepare Documents for LLMs, Embeddings, and RAG

Understand chunking vs tokenization in LLM and RAG systems. Learn chunk size, overlap, semantic splitting, token-based chunking, retrieval tradeoffs, and how to evaluate the best strategy.
Retrieval units vs model units
Chunk size & overlap
Evaluate, don't default

Chunking and tokenization both break information into smaller units, but they solve different problems.

Tokenization converts text into the units a language model processes.

Chunking divides documents into retrievable units that can be embedded, indexed, searched, and supplied as context to an LLM.

A useful mental model is:

Document → chunks → tokens → embeddings / retrieval → LLM context

Chunking determines what information is grouped together. Tokenization determines how the model represents and counts that text internally.

Chunking vs tokenization

DimensionChunkingTokenization
PurposeCreate retrieval unitsCreate model-processing units
Typical unitsection, paragraph, semantic blocktoken/subword
Controlled by applicationusuallyusually model/tokenizer dependent
Affects retrieval precisionstronglyindirectly
Affects context-window useyesdirectly

A chunk usually contains many tokens.

Why tokenization matters

Language models do not process text exactly as people see words. A tokenizer can split text into words, subwords, punctuation, symbols, or code fragments.

That means:

characters ≠ words ≠ tokens

Token counts matter because context limits and usage are token-based.

Why chunking matters in RAG

Imagine a 200-page operations manual. Embedding the whole document as one vector makes retrieval too coarse. Instead, split the document into smaller units and attach useful metadata such as heading, page, effective date, source URL, and access level.

Chunks that are too large

Oversized chunks reduce specificity, send irrelevant content to the LLM, increase token cost, and can introduce competing facts.

Chunks that are too small

Tiny chunks can lose headings, exceptions, definitions, table labels, and surrounding reasoning.

The objective is:

the smallest chunk that preserves enough information to answer the intended class of questions correctly.

Common chunking strategies

Fixed-size chunking

Simple and deterministic, but may split ideas at awkward boundaries.

Recursive chunking

Uses natural boundaries such as sections and paragraphs before falling back to smaller units. This is a strong general baseline.

Token-based chunking

Uses a token maximum to control embedding and LLM context limits.

Sentence or paragraph chunking

Preserves more semantic coherence than arbitrary character boundaries.

Structure-aware chunking

Uses HTML headings, Markdown sections, JSON objects, code functions, legal clauses, or other source structure.

Semantic chunking

Splits when meaning or topic changes rather than only at fixed length.

LLM-based chunking

Uses a model to identify logical units or rewrite material into self-contained propositions. It can be powerful but adds cost, latency, and governance considerations.

Chunk overlap

Overlap repeats part of one chunk in the next. It can preserve evidence that crosses a boundary, but it also creates duplicate embeddings, larger indexes, redundant retrieval, and extra context tokens.

Treat overlap as a measured parameter, not an automatic default.

Parent-child chunking

A useful pattern is to index smaller child chunks for retrieval precision while preserving a reference to a larger parent section for reasoning context.

This gives:

small retrieval unit + larger reasoning context

Tables need special handling

A row without its headers may be meaningless. Preserve structural context when turning tables into retrievable text.

PDFs need special handling

PDF extraction can introduce repeated headers, broken reading order, OCR errors, and damaged tables. Separate extraction quality from chunking quality.

Metadata and lineage

A production retrieval unit is better thought of as:

chunk text + metadata + lineage

Metadata can enable filters, security, freshness controls, source attribution, reranking, and traceability.

Chunking is part of context engineering

The information an LLM sees depends on parsing, chunking, embeddings, metadata, retrieval, reranking, context assembly, and prompting.

Changing chunking changes the model's effective information environment.

How to choose chunk size

Do not use one default because a framework tutorial chose it. Build an evaluation set and compare several strategies, for example:

  • 250 tokens,
  • 500 tokens,
  • 1,000 tokens,
  • section-aware splitting,
  • semantic splitting.

Then run the same questions against each configuration.

Metrics

Useful retrieval metrics include Recall@K, Precision@K, MRR, and NDCG.

Then measure answer correctness, faithfulness, context efficiency, latency, embedding/index cost, and downstream LLM cost.

Common mistakes

One chunk size for every content type.

Optimizing only for maximum context length.

Ignoring headings.

Excessive overlap.

Losing source lineage.

Failing to re-chunk when sources change.

Measuring only final answers.

Multilingual data

Chunking rules that work well for English may not generalize to all writing systems. Test by language.

Long-context LLMs do not eliminate chunking

Retrieval still provides access control, provenance, freshness, lower input cost, targeted evidence, and scalability across large knowledge collections.

Practical enterprise default

Preserve headings.

Use recursive structure-aware splitting.

Enforce a token-aware maximum.

Use modest overlap.

Preserve metadata and lineage.

Retrieve a small Top-K.

Measure Recall@K and answer quality.

Then add semantic chunking, parent-child retrieval, late chunking, proposition extraction, or reranking only when evaluation identifies a need.

Chunking should be a knob

Treat chunk strategy, size, overlap, separator hierarchy, metadata inclusion, parent expansion, retrieval K, reranking K, and embedding model as measurable variables.

Bottom line

Tokenization converts text into units processed by the model. Chunking converts documents into units retrieved by the information system.

Do not ask only "What chunk size should we use?" Ask "Which chunking configuration produces the best outcomes for our documents and queries?"

KNOBS

Treat chunk size, overlap, embedding model, Top-K, and reranking as measurable configuration variables rather than hidden defaults.

Explore KNOBS →
FAQ

Frequently asked questions

Short answers to the questions teams ask most often about chunking vs tokenization.

What is the difference between chunking and tokenization?

Chunking creates retrieval units; tokenization creates model-processing units. One chunk generally contains many tokens.

What is the best chunk size for RAG?

There is no universal optimum. Test multiple sizes and strategies against retrieval and answer-quality metrics.

Do long-context LLMs eliminate chunking?

No. Retrieval still helps with permissions, provenance, freshness, cost, and scale.