AI MuseumELEL

Exhibit 6.4

4

Instruction tuning

With examples of instructions and high-quality answers, the model is trained to follow requests and behave more like an assistant.

Instruction tuning
AI-generated illustration

Why it is in the museum

A modern language model does not appear fully formed. It passes through data collection, training, adaptation, evaluation and, finally, answer generation.

What supports this exhibit

Primary research paper

Supervised instruction tuning with human-written demonstrations before RLHF.

Main source: Ouyang, L. et al. (2022), InstructGPT.

Open the source

What to keep in mind: Primary source for one important family of techniques.

Ask the exhibit

This small guide uses only the information documented on this exhibit page. If your question goes beyond that evidence, it will say so instead of inventing an answer.

Start with one of the suggested questions above, or type your own.