Back to archive
#ai#llm#glossary#aigen

Autoregressive Language Modeling

You start a sentence: “Today I drink…”. After appending “tea”, possible continuations differ from those after “water”. The model must account for the text already chosen at every next step, rather than guess all positions independently.

Autoregressive Language Modeling describes text through the probabilities of successive tokens, pieces of text that depend on earlier tokens. The model determines probabilities of possible continuations, and a separate rule chooses what to append.

After choosing “tea”, that fragment becomes part of the next step's context. The probability of the entire sequence combines the probabilities of successive choices. During training, the text is known, so predictions can be assessed against the actual following fragments.

“Autoregressive” specifies a way of modeling, rather than the name of one architecture. It also does not mean the most probable continuation is true. Causal Masking enables training many positions at once without looking at their future answers.

Mechanism and details

For a sequence of fixed length, the expression is:

pθ(x1,…,xT)=∏t=1Tpθ(xt∣x<t)p_\theta(x_1,\ldots,x_T)=\prod_{t=1}^{T}p_\theta(x_t\mid x_{<t})

xtx_t is the token at position tt, x<tx_{<t} is the preceding prefix, and θ\theta denotes the model's parameters. The first factor has no preceding text tokens; an implementation may use a beginning marker. A model with a limited context window uses only the available suffix of the prefix. This distribution and a neural next-word prediction model are described by Bengio et al., A Neural Probabilistic Language Model, §1.2 and §2.

Training and generation

An original example: after the text “Today I drink,” the model assigns probabilities to possible continuations. When the generation procedure selects a token corresponding to “tea,” it becomes part of the prefix for the next step. The distribution itself does not determine whether to select the highest probability or sample.

During training, the correct sequence is known. The optimizer increases the log-probability of its successive tokens. The original GPT, §3.1, applies this objective to text tokens using a Transformer decoder.

Causal Masking, together with shifted targets, makes it possible to calculate predictions for many positions in parallel without revealing the answers. Vaswani et al., §3.1, describe this for the decoder. Generated tokens still depend sequentially on earlier choices.

This is a way of modeling a distribution, rather than the name of an architecture: it existed before Transformer. In a typical text-generating model, LM head supplies logits from which Softmax derives the next-token distribution.