08 · LLMs · 6 min read · Interactive · updated
What is tokenization in LLMs and how does the BPE algorithm work?
In short
Tokenization splits text into tokens, chunks of words that a model turns into numbers. BPE builds the vocabulary by merging the most frequent symbol pairs.
What it is
Tokenization is the conversion of text into a sequence of tokens — units from a fixed vocabulary, each with its own ID number. A language model never sees letters or words, only these numbers. Byte-Pair Encoding (BPE) is the most popular algorithm for building such a vocabulary: it starts from single characters (or bytes) and repeatedly merges the most frequent adjacent pair into a new symbol.
Why not words? A vocabulary of whole words would have to be gigantic and would still miss new names, typos and inflected forms — in a highly inflected language such as Polish, a single word has a dozen or more forms. Why not letters? Sequences would be very long, and the model would have to assemble meaning from letters from scratch. Subword tokens are the compromise: frequent words become a single token, rare ones break down into several familiar pieces.
The vocabulary of a typical modern LLM has from tens of thousands to a few hundred thousand tokens. In English text one token averages about three to four characters; in Polish usually fewer, because vocabularies are built mainly on English data.
Mechanism — why it works this way
Training the tokenizer. BPE in the version of Sennrich et al. (2016) works on a corpus with word counts. Each word is written as a sequence of characters followed by an end-of-word symbol. Then, in a loop: (1) count all pairs of adjacent symbols, weighting by the number of occurrences of each word; (2) replace the most frequent pair with a new symbol and append the merge rule to a list; (3) repeat until the vocabulary reaches the target size. The result is an ordered list of merges.
Tokenizing new text. The text is split into characters and the merges are applied in the same order in which they were learned. Because you can always fall back to single characters, every word can be written down — there is no "unknown word" problem. The byte-level version (GPT-2) starts from the 256 possible UTF-8 bytes, so it handles any text, including emoji and characters from other scripts.
Why frequency is a good heuristic. Merging the most frequent pairs is greedy compression: each merge shortens the encoded corpus the most. Shorter sequences mean more content fits in the context window, and the model performs fewer steps for the same text. Along the way, frequent stems and endings get their own tokens, which helps the model learn morphology — although BPE splits do not coincide with morpheme boundaries.
Consequences visible in model behaviour. The model does not "see" the letters inside a token, which is why tasks such as counting the letters in a word or reversing it go badly. Numbers are often split into irregular chunks, which makes arithmetic harder. Languages that are under-represented in the tokenizer's training data need more tokens for the same sentence: they pay more, wait longer and fill the context faster (Petrov et al., 2023). Polish letters with diacritics take two bytes each in UTF-8, so in a byte-level tokenizer they less often merge into long tokens.
Variants. WordPiece (BERT) chooses the merge that most increases the likelihood of the data, rather than raw frequency. The Unigram LM (SentencePiece, Kudo and Richardson, 2018) works the other way round: it starts with a large vocabulary and removes the least useful tokens. The principle stays the same: a subword vocabulary fitted to the statistics of the corpus.
By example
Take a mini-corpus of Polish words: "dom" (house) ×6, "domu" (of the house) ×3, "domy" (houses) ×2, "kot" (cat) ×4, "kota" (of the cat) ×1, "kotu" (to the cat) ×2 (the symbol _ marks the end of a word). At the start the encoding takes 80 symbols. The pairs "d o" and "o m" each occur 11 times — we break the tie in favour of the first. The merges that follow: (1) d+o → "do" (11), (2) do+m → "dom" (11), (3) k+o → "ko" (7, tied with "o t"), (4) ko+t → "kot" (7), (5) dom+_ → "dom_" (6), (6) u+_ → "u_" (5, since "domu" and "kotu" give 3 + 2 together), (7) kot+_ → "kot_" (4).
After seven merges, "dom" is a single token, "domu" is dom·u_, "kotu" is kot·u_, and the whole corpus takes 29 symbols instead of 80. The algorithm has separated the stems and the case ending "-u" on its own, although it knows nothing about grammar. A new word, "domek" (little house), which was not in the corpus, will be encoded as dom·e·k·_ — a stem plus letters. It works the same way at scale: GPT-2 has a vocabulary of 50,257 tokens, i.e. 256 base bytes, 50,000 learned merges and one special token (Radford et al., 2019).
In practice
- To count tokens, use the tokenizer of the specific model:
AutoTokenizer.from_pretrained(...)intransformers,tiktokenorsentencepiece. Different models split the same text differently. - You can train your own tokenizer, e.g. with the
tokenizerslibrary (BPE,BpeTrainer) — it is worth it only when training a model from scratch. - Context limits and prices are given in tokens. Text in languages other than English, Polish included, usually uses noticeably more of them than English text with the same content — measure it on your own data.
- A space is usually part of the token (" dom" and "dom" are different tokens), which matters when building prompts and parsing output.
- Common mistake: adding new tokens to the vocabulary without training their embeddings — the model treats them as noise.
Frequently asked questions
- How many words is one token?
- It depends on the language and the tokenizer. For English the popular rule of thumb is about 0.75 words per token; for Polish there are more tokens per word. The only reliable way is to count with the given model's tokenizer.
- Why do LLMs make mistakes counting letters?
- Because they receive tokens, not letters. The word "strawberry" may be two or three tokens, and the model would have to know from memory which letters make up each of them. Spelling the word out letter by letter in the prompt helps.
- Can the tokenizer of a finished model be changed?
- Not without cost. The embeddings and the output layer are tied to specific token IDs, so changing the vocabulary requires further training. More often a few special tokens are added and the model is fine-tuned.
Sources
- Sennrich R., Haddow B., Birch A., 2016, "Neural Machine Translation of Rare Words with Subword Units", ACL 2016.
- Radford A. et al., 2019, "Language Models are Unsupervised Multitask Learners", OpenAI technical report.
- Kudo T., Richardson J., 2018, "SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing", EMNLP 2018 (System Demonstrations).
- Petrov A. et al., 2023, "Language Model Tokenizers Introduce Unfairness Between Languages", NeurIPS 2023.
- Jurafsky D., Martin J. H., "Speech and Language Processing", 3rd ed. (online draft), ch. 2 (words and tokens).