Room 6, the production line

How an LLM is made today

A modern language model is not written line by line by programmers. It passes through a sequence of stages, more like a production line, involving data, compute, training, evaluation and deployment. Each stage adds capabilities and introduces trade-offs.

Open the interactive room →

Exhibits in this room

1

Data collection

Huge quantities of text are gathered from sources such as websites, books and code. Selection and cleaning influence what a model can learn, how well it performs and which biases it may reproduce.

2

Tokenization

Text is split into smaller units called tokens, which may be whole words or pieces of words. The model processes numerical identifiers rather than written characters directly.

3

Pretraining

The model repeatedly learns to predict the next token across enormous text corpora and large amounts of computing. This stage gives it broad language ability and general patterns, but not necessarily assistant-like behaviour.

4

Instruction tuning

With examples of instructions and high-quality answers, the model is trained to follow requests and behave more like an assistant.

5

Alignment

People, and sometimes other models, evaluate answers. Training then pushes the model toward responses judged more helpful, honest or safe. A well-known approach is reinforcement learning from human feedback (RLHF).

6

Testing and safety

Before release, models are evaluated for errors, harmful behaviour and misuse. Testing and monitoring can continue after deployment.

7

Use

When you type a prompt, a basic language model does not retrieve a ready-made answer from a database. It generates a response token by token, although some modern systems can also call search engines, tools or external knowledge sources.