Masked Language Modeling
You want to teach the model to use the entire sentence. You hide a fragment in “Alice has a [MASK]” or “Alice [MASK] a cat” and ask it to reconstruct the original. Information on both sides of the gap may be needed for the answer.
Masked Language Modeling (MLM) is the task of reconstructing selected tokens after corrupting the text. A token is a piece of text. The model compares predictions at selected positions with the original fragments and changes its parameters on that basis.
In [A, MASK, C], the correct middle answer remains B from the original [A, B, C]. The model can use both A and C. BERT used this training approach with specified proportions of masking and other changes.
This meaning of “masking” differs from Causal Masking, which blocks access to future positions. BERT's recipe is not compulsory for every MLM, and reconstructing gaps is not yet a rule for writing a complete chat response.
Mechanism and details
In the BERT paper, §3.1, the authors randomly selected 15% of WordPiece token positions. Of the selected positions, they replaced 80% with [MASK], 10% with a random token, and left 10% unchanged. These proportions describe BERT's specific recipe; the concept of MLM does not require these particular values.
What the model must reconstruct
An original illustration: we have three illustrative tokens [A, B, C]. After changing the input to [A, MASK, C], the target at the middle position is still B. The model can use information from A and C. If we insert a random token D instead of the mask, the correct answer remains the original B.
The MLM loss term, based on Cross-entropy, is computed at selected positions against their original tokens. It also includes selected positions that were ultimately left unchanged. Appendix C.2 compares different corruption proportions; this treatment reduces the difference between inputs during pre-training and later tasks without [MASK].
Perturbing tokens should not be equated with Causal Masking. There, a mask limits attention's access to future positions. In MLM as described for BERT, access to context on both sides is part of the task. The original BERT also had a separate Next Sentence Prediction objective, so MLM does not describe all of its training.