Language models never see letters — they see tokens, and every bill is counted in them. This runs a real byte-pair encoder in your browser. Type anything, then drag the merge slider to watch the vocabulary get built one merge at a time.
merges applied 800 · vocabulary 1,056 tokens · this text 0 tokens
At zero merges every UTF-8 byte is its own token. Each merge glues the most frequent adjacent pair in the training corpus into one new symbol — the first few buy enormous savings, the eight-hundredth barely registers. That curve is why production vocabularies stop around 100k.
⟨F0⟩ is a raw byte with no vocabulary entry of its own — emoji and non-Latin
scripts land there. · is a space, ⏎ a newline.
Every step is one entry from the embedded merge table, applied lowest rank first — exactly the algorithm, just with a much smaller table.
| Example model | $ / 1M in | this text | ×10,000 runs |
|---|
Billed on 0 input tokens. Rates are editable examples (public list prices, August 2026) — always check your provider's current pricing.
Same mechanism, tiny scale. This is genuine byte-pair encoding: text is split by a GPT-2-style pre-tokenizer regex, each piece is converted to UTF-8 bytes, and merges from the embedded table are applied in the order they were learned. The merge table was trained offline on a few megabytes of English documentation, prose and Python source.
What's different. 800 merges give a 1,056-token vocabulary. Production tokenizers carry 100,000–200,000 tokens learned from terabytes of multilingual text. So counts here run roughly 1.8–2× higher than a real tokenizer on the same English text, and far higher on other scripts. There are also no special tokens — no BOS/EOS, no chat template.
What still transfers. Whitespace belongs to the following word, so " the"
and "the" are different tokens. Emoji are four UTF-8 bytes and shatter into byte tokens. Rare
words fragment while common ones survive whole. Non-English scripts pay a steep tax. Long digit runs are
expensive. Those effects are all real, and they are why token counts never match character counts.