Back to archive
#ai#llm#glossary#aigen

Transformer

You want to translate “The train will arrive tomorrow”. Translating each word separately is insufficient: the correct form depends on other words and their order. The model needs a way to combine information from text, even when important fragments are far apart.

A Transformer is a neural network architecture, a program learning computations from examples. It receives text already divided into tokens — pieces of text — and creates their descriptions as lists of numbers. A separate program called a tokenizer performs the division. Successive blocks mix information from available positions and transform each position's description. The mechanism combining context is called Self-attention.

In our sentence, the description of “arrive” can account for “train” and “tomorrow”. This is not a hand-written grammar rule: the mixing is learned. A separate mechanism, such as Positional Encoding, supplies order information.

Much of this computation can run in parallel during training. This does not mean a new response appears all at once: many models still append it piece by piece. Transformer describes a model's construction, rather than guaranteeing truthfulness or understanding of text.

Mechanism and details

At the input, token IDs are converted into vectors. Positional Encoding supplies information about their order. Within a block, Multi-Head Attention combines information from available positions, and the feed-forward network transforms the result separately for each position. Residual connections and normalization help pass the signal through successive blocks. These components can be seen in the TensorFlow implementation.

Encoder and decoder

The original model in Vaswani et al., Attention Is All You Need, §3 had an encoder and a decoder. The encoder created a representation of the input sentence, and the decoder used it when generating a translation.

For example, when translating “The train will arrive tomorrow,” the decoder produces a probability distribution for the next token based on the input and the already available part of the response. Calculations for many positions can be performed in parallel during training; generating a new response can still proceed token by token.

The name Transformer does not specify a single model size or training objective. Encoder-only and decoder-only variants exist, so details of the original arrangement should not be attributed to every LLM. TensorFlow discusses these variants at the end of the guide.