Skip to content
Ishaan Reddy

Cookbook · LLMs

Tokenization

5 mintokenizationbpellms

The idea, in one analogy

Imagine you had to communicate using only a fixed deck of a few thousand flashcards, each with a word or word-fragment printed on it. "Unbelievable" might not have its own card, but "un", "believ", and "able" do, so you hold up three cards instead of one. A tokenizer is that deck: a fixed vocabulary of chunks, and a rule for splitting any input text into a sequence of them.

Why not just use whole words, or single letters?

Single characters: tiny vocabulary, but every sentence becomes a very long sequence of tokens, and the model has to rebuild "word" as a concept from scratch.

Whole words: short sequences, but the vocabulary needs to be enormous to cover every word form in every language (and it still fails the moment it sees a word it's never seen: a new brand name, a typo, a rare technical term).

Byte-Pair Encoding (BPE), used by nearly every modern LLM tokenizer, is the middle ground: start with individual characters (or bytes), then repeatedly merge the most frequent adjacent pair into a new token, thousands of times, until you hit a target vocabulary size. Common words end up as single tokens ("the" → 1 token); rare or unfamiliar text falls back to smaller pieces, down to individual bytes in the worst case, so nothing is ever "unencodable."

A worked example: training BPE on "low lower lowest"

Start with every character its own token (</w> marks a word boundary):

Step 0:  l o w </w>  |  l o w e r </w>  |  l o w e s t </w>

Count every adjacent pair across the whole tiny corpus: (l,o) appears 3 times, (o,w) appears 3 times, (w,e) appears 2 times, and several others appear once.

(l,o) and (o,w) are tied at 3, so pick one (implementations break ties by scan order; here, (l,o)) and merge it everywhere:

Step 1:  lo w </w>  |  lo w e r </w>  |  lo w e s t </w>

Recount. Now (lo,w) is the most frequent pair (3 occurrences), so merge that next:

Step 2:  low </w>  |  low e r </w>  |  low e s t </w>

Recount again. (e,r) and (e,s) each appear once; suppose (e,r) wins this round and gets merged:

Step 3:  low </w>  |  low er </w>  |  low e s t </w>

A few more rounds like this and the corpus's most common substrings, low, er, est, end up as single tokens, while everything else stays as smaller pieces. Run this same process over billions of words instead of three, and the merge list that falls out is a real BPE vocabulary.

The decision that matters most: whose tokenizer?

Using a widely-used pretrained tokenizer (GPT-2's, or a modern one like tiktoken's cl100k_base) means compatibility with existing tooling and no training step, but it was built on that project's data distribution, not yours. Training your own BPE tokenizer on your own corpus gets you better compression (fewer tokens per unit of text, which directly reduces training and inference cost) on your specific domain, at the cost of an extra pipeline step and drop-in compatibility with anything expecting the original vocabulary.

A concrete, measured example: a domain-specific tokenizer can badly fragment vocabulary it wasn't trained on. In practice, a general-web 32K BPE tokenizer split the word "photosynthesis" into 4 pieces; a science-specific tokenizer trained on the right corpus would likely keep it as one or two. That fragmentation isn't just cosmetic: more tokens per concept means more steps for the model to piece the idea back together, and less context fits in a fixed window.

Vocabulary choice also changes how many tokens the same sentence costs. Take "The quick brown fox jumps over the lazy dog.":

TokenizerToken countBreakdown
Character-level45T, h, e, , q, u, i, c, k, ... (one token per character)
GPT-2 BPE (50,257 vocab)10The, quick, brown, fox, jumps, over, the, lazy, dog, .
A well-trained 32K BPE tokenizer8The, quick, brown, fox, jumps, over, the, lazy dog.

Same sentence, 5.6x fewer tokens between the least and most compressed version. Every one of those saved tokens is a step the model doesn't have to spend attention on, and a slot in the context window that goes to something else instead.

Where to look further