Chunking and tokenization both break information into smaller units, but they solve different problems.
Tokenization converts text into the units a language model processes.
Chunking divides documents into retrievable units that can be embedded, indexed, searched, and supplied as context to an LLM.
A useful mental model is:
Chunking determines what information is grouped together. Tokenization determines how the model represents and counts that text internally.
Chunking vs tokenization
| Dimension | Chunking | Tokenization |
|---|---|---|
| Purpose | Create retrieval units | Create model-processing units |
| Typical unit | section, paragraph, semantic block | token/subword |
| Controlled by application | usually | usually model/tokenizer dependent |
| Affects retrieval precision | strongly | indirectly |
| Affects context-window use | yes | directly |
A chunk usually contains many tokens.
Why tokenization matters
Language models do not process text exactly as people see words. A tokenizer can split text into words, subwords, punctuation, symbols, or code fragments.
That means:
Token counts matter because context limits and usage are token-based.
Why chunking matters in RAG
Imagine a 200-page operations manual. Embedding the whole document as one vector makes retrieval too coarse. Instead, split the document into smaller units and attach useful metadata such as heading, page, effective date, source URL, and access level.
Chunks that are too large
Oversized chunks reduce specificity, send irrelevant content to the LLM, increase token cost, and can introduce competing facts.
Chunks that are too small
Tiny chunks can lose headings, exceptions, definitions, table labels, and surrounding reasoning.
The objective is:
Common chunking strategies
Fixed-size chunking
Simple and deterministic, but may split ideas at awkward boundaries.
Recursive chunking
Uses natural boundaries such as sections and paragraphs before falling back to smaller units. This is a strong general baseline.
Token-based chunking
Uses a token maximum to control embedding and LLM context limits.
Sentence or paragraph chunking
Preserves more semantic coherence than arbitrary character boundaries.
Structure-aware chunking
Uses HTML headings, Markdown sections, JSON objects, code functions, legal clauses, or other source structure.
Semantic chunking
Splits when meaning or topic changes rather than only at fixed length.
LLM-based chunking
Uses a model to identify logical units or rewrite material into self-contained propositions. It can be powerful but adds cost, latency, and governance considerations.
Chunk overlap
Overlap repeats part of one chunk in the next. It can preserve evidence that crosses a boundary, but it also creates duplicate embeddings, larger indexes, redundant retrieval, and extra context tokens.
Treat overlap as a measured parameter, not an automatic default.
Parent-child chunking
A useful pattern is to index smaller child chunks for retrieval precision while preserving a reference to a larger parent section for reasoning context.
This gives:
Tables need special handling
A row without its headers may be meaningless. Preserve structural context when turning tables into retrievable text.
PDFs need special handling
PDF extraction can introduce repeated headers, broken reading order, OCR errors, and damaged tables. Separate extraction quality from chunking quality.
Metadata and lineage
A production retrieval unit is better thought of as:
Metadata can enable filters, security, freshness controls, source attribution, reranking, and traceability.
Chunking is part of context engineering
The information an LLM sees depends on parsing, chunking, embeddings, metadata, retrieval, reranking, context assembly, and prompting.
Changing chunking changes the model's effective information environment.
How to choose chunk size
Do not use one default because a framework tutorial chose it. Build an evaluation set and compare several strategies, for example:
- 250 tokens,
- 500 tokens,
- 1,000 tokens,
- section-aware splitting,
- semantic splitting.
Then run the same questions against each configuration.
Metrics
Useful retrieval metrics include Recall@K, Precision@K, MRR, and NDCG.
Then measure answer correctness, faithfulness, context efficiency, latency, embedding/index cost, and downstream LLM cost.
Common mistakes
One chunk size for every content type.
Optimizing only for maximum context length.
Ignoring headings.
Excessive overlap.
Losing source lineage.
Failing to re-chunk when sources change.
Measuring only final answers.
Multilingual data
Chunking rules that work well for English may not generalize to all writing systems. Test by language.
Long-context LLMs do not eliminate chunking
Retrieval still provides access control, provenance, freshness, lower input cost, targeted evidence, and scalability across large knowledge collections.
Practical enterprise default
Preserve headings.
Use recursive structure-aware splitting.
Enforce a token-aware maximum.
Use modest overlap.
Preserve metadata and lineage.
Retrieve a small Top-K.
Measure Recall@K and answer quality.
Then add semantic chunking, parent-child retrieval, late chunking, proposition extraction, or reranking only when evaluation identifies a need.
Chunking should be a knob
Treat chunk strategy, size, overlap, separator hierarchy, metadata inclusion, parent expansion, retrieval K, reranking K, and embedding model as measurable variables.
Bottom line
Tokenization converts text into units processed by the model. Chunking converts documents into units retrieved by the information system.