Back to archive
#ai#llm#glossary#aigen

Causal Masking

You are teaching the model to append “cat” after “Alice has a”. The word “cat” already exists in the complete example. If the model could look at it while predicting, the task would appear easy but would not resemble creating a new response.

Causal Masking blocks information flow from later text positions to earlier ones. A token is a piece of text. When the position containing “a” helps predict the next token, it can use “Alice”, “has”, and “a”, but the future “cat” cannot affect its result.

This blocks access, rather than merely reducing the importance of a future fragment. Its contribution to information mixing must be zero. During training, many positions can therefore be processed simultaneously while preserving prediction conditions based solely on the available past.

The mask does not explain word meanings or replace the positional signal. The table and experiment below show which connections remain allowed.

Mechanism and details

The mask controls which connections are allowed, not word meanings. A rejected position is not merely given a slightly smaller weight: its contribution to the attention result must be zero. The TensorFlow implementation shows this condition and compares results for full and shortened sequences.

Which positions are available?

For three positions, we get:

Query positionKey 1Key 2Key 3
1yesnono
2yesyesno
3yesyesyes

When training next-token prediction, input [start, A, B] can be paired with targets [A, B, C]. The second position sees start and A when predicting B. Access to the current input position therefore does not reveal the answer. This illustrates shifting targets by one step.

In Scaled Dot-Product Attention, forbidden comparison scores are mathematically replaced by negative infinity before softmax. The mask and shift are described by Vaswani et al., §3.1 and §3.2.3.

Positional Encoding supplies order information but does not itself block access. A padding mask solves yet another problem: it ignores artificial sequence padding, regardless of whether it lies in the future.

See also: Sequence Packing — when combining independent examples, blocking future positions alone does not preserve their boundaries.