Loss masking
You are training an agent on records of conversations with tools. A record contains its commands and the tools' responses. You want to assess selected parts while still letting it read the rest as context for its next action.
Loss masking determines which positions contribute to the loss, the numerical measure of error during learning. A token — a piece of text — may remain visible in the input even though the model is not directly assessed for predicting it.
When training only on the agent's responses, the user's text and the tool's result help it choose the next step, but receive no separate prediction penalty. The model therefore does not need to learn to generate every part of the entire record.
This differs from Causal Masking, which restricts information flow between positions. Having no loss term for a token need not mean having no effect on training: its computations may influence later positions that are assessed. “Don't Mask the Environment”, §2, examines the consequences of choosing which parts to assess.
See also: Sequence Packing shows why combining examples requires separate control over access to context and the positions assessed by the loss.
