AI MuseumELEL

Exhibit 6.1

1

Data collection

Huge quantities of text are gathered from sources such as websites, books and code. Selection and cleaning influence what a model can learn, how well it performs and which biases it may reproduce.

Data collection
AI-generated illustration

Why it is in the museum

A modern language model does not appear fully formed. It passes through data collection, training, adaptation, evaluation and, finally, answer generation.

What supports this exhibit

Curatorial synthesis of technical literature

That data selection and curation influence model knowledge, quality and bias.

Main source: Technical literature on pretraining, data curation and scaling of modern LLMs.

The source is listed here, but no direct link is currently available in the register.

What to keep in mind: General process; details vary by model.

Ask the exhibit

This small guide uses only the information documented on this exhibit page. If your question goes beyond that evidence, it will say so instead of inventing an answer.

Start with one of the suggested questions above, or type your own.