Pre-training
You want the model to recognize relationships in language before adapting it to a specific job, such as classifying customer messages. Manually preparing answers for all the text would be impractical. The text itself, however, can supply training tasks.
Pre-training is the stage of developing parameters that provide the starting point for further model use or adaptation. Parameters are the numbers controlling its computations. In language models, the signal often comes from a large collection of texts.
“Alice has a cat” can supply a task of predicting the next fragment after “Alice has a” or reconstructing the hidden word in “Alice [MASK] a cat”. Both exercises use existing text, although they train with different access to context.
The stage's name does not impose one task or architecture. After pre-training, a model need not yet follow user instructions well. Fine-tuning can later adapt it to chosen examples and behaviors.
Mechanism and details
The name defines the stage's role in the learning process. It does not impose a single architecture or objective function.
The same stage, different tasks
In the original GPT, the model learns to predict the next token from preceding ones. This is Autoregressive Language Modeling. The authors describe this stage before adaptation to labeled tasks in Improving Language Understanding by Generative Pre-Training, §3.1–3.2.
The original BERT uses a different set of objectives: it reconstructs selected tokens after perturbing the input and identifies whether two passages followed one another in the corpus. When reconstructing a token, it can use context on both sides. These are specific tasks from the BERT paper, §3.1, not a mandatory set for every model.
An original example of the difference: the sequence A B C can be used to prepare a task predicting C after A B, or reconstructing B from A [MASK] C. Both signals come from the available text, but they require a different flow of information. The absence of a manually assigned label does not mean the absence of a training objective.
What remains after training
The saved model stores learned parameters. These include the Token Embedding table, which in GPT is trained together with the rest of the network. The next stage, Fine-tuning, starts from parameters that have already been learned and adapts the model to a specific use.
The word “pre-trained” alone does not say what data the model trained on, what it optimized, or how well it will perform a particular task. Those details must be read from the description of the model.