AI MuseumELEL

Exhibit 6.2

2

Tokenization

Text is split into smaller units called tokens, which may be whole words or pieces of words. The model processes numerical identifiers rather than written characters directly.

Tokenization
AI-generated illustration

Why it is in the museum

A modern language model does not appear fully formed. It passes through data collection, training, adaptation, evaluation and, finally, answer generation.

What supports this exhibit

Primary research paper

The use of subword units such as BPE to represent rare or previously unseen words.

Main source: Sennrich, Haddow & Birch (2016), ACL, subword units.

Open the source

What to keep in mind: Primary source; exact tokenizers differ across models.

Ask the exhibit

This small guide uses only the information documented on this exhibit page. If your question goes beyond that evidence, it will say so instead of inventing an answer.

Start with one of the suggested questions above, or type your own.