Tokenizer Visualizer

Language models never see letters — they see tokens, and every bill is counted in them. This runs a real byte-pair encoder in your browser. Type anything, then drag the merge slider to watch the vocabulary get built one merge at a time.

800 merges 1,056-token vocab trained offline on English docs + code approximation, not a real LLM tokenizer
tokens
0
characters
0
words
0
utf-8 bytes
0
chars / token
0

Input text

—
samples:

Vocabulary budget

0 merges800 merges

merges applied 800 · vocabulary 1,056 tokens · this text 0 tokens

At zero merges every UTF-8 byte is its own token. Each merge glues the most frequent adjacent pair in the training corpus into one new symbol — the first few buy enormous savings, the eight-hundredth barely registers. That curve is why production vocabularies stop around 100k.

Tokens

—

⟨F0⟩ is a raw byte with no vocabulary entry of its own — emoji and non-Latin scripts land there. · is a space, ⏎ a newline.

Merge trace

click any chip

Every step is one entry from the embedded merge table, applied lowest rank first — exactly the algorithm, just with a much smaller table.

Cost estimate

count basis:
Example model$ / 1M inthis text×10,000 runs

Billed on 0 input tokens. Rates are editable examples (public list prices, August 2026) — always check your provider's current pricing.

▸ token IDs for this text
▸ how this differs from a real LLM tokenizer

Same mechanism, tiny scale. This is genuine byte-pair encoding: text is split by a GPT-2-style pre-tokenizer regex, each piece is converted to UTF-8 bytes, and merges from the embedded table are applied in the order they were learned. The merge table was trained offline on a few megabytes of English documentation, prose and Python source.

What's different. 800 merges give a 1,056-token vocabulary. Production tokenizers carry 100,000–200,000 tokens learned from terabytes of multilingual text. So counts here run roughly 1.8–2× higher than a real tokenizer on the same English text, and far higher on other scripts. There are also no special tokens — no BOS/EOS, no chat template.

What still transfers. Whitespace belongs to the following word, so " the" and "the" are different tokens. Emoji are four UTF-8 bytes and shatter into byte tokens. Rare words fragment while common ones survive whole. Non-English scripts pay a steep tax. Long digit runs are expensive. Those effects are all real, and they are why token counts never match character counts.