A model’s training examples reflect choices about collection, selection, and representation. For a supervised task, labels also define the distinctions the model is being asked to learn.
Consider whether the examples cover the situations in which the system will be used. A collection can be large while leaving important cases underrepresented. Inconsistent labels can make the intended task unclear even when the inputs themselves are plentiful.
When evaluating a system, ask how the task and data relate to the real use case. The amount of data is one part of the story; its relevance, quality, and coverage are others. Those questions help connect model performance with the work it is expected to support.
A small working example.
Imagine teaching a formatting task through several similar notes. Add a deliberately different note to examine which relationships the examples actually communicate.
A note to keep beside it.
Treat an example as part of the instruction. Its omissions and boundaries can influence a result as much as the words that describe the task.
- Ask which situations the examples cover.
- Look for a clear labeling rule.
- Connect the training task with the intended use.
Follow a related question
List the assumptions behind a prediction.
A thoughtful view of what comes nextWrite the question beside the idea.
Capture the question firstKeep learning
Related background to continue exploring this subject.
Google: an introduction to language models NIST: AI risk management framework

