Exhibit 6.3
3Pretraining
The model repeatedly learns to predict the next token across enormous text corpora and large amounts of computing. This stage gives it broad language ability and general patterns, but not necessarily assistant-like behaviour.
Why it is in the museum
A modern language model does not appear fully formed. It passes through data collection, training, adaptation, evaluation and, finally, answer generation.
What supports this exhibit
Curatorial synthesis of technical literature
Training with next-token prediction over very large text corpora.
Main source: Research literature on autoregressive language modelling and large-scale pretraining.
The source is listed here, but no direct link is currently available in the register.
What to keep in mind: Generalised description; not every LLM uses exactly the same training procedure.
Ask the exhibit
This small guide uses only the information documented on this exhibit page. If your question goes beyond that evidence, it will say so instead of inventing an answer.
Start with one of the suggested questions above, or type your own.