Token Embedding
The computer receives a number identifying a text fragment. The number 17 itself does not tell it what the fragment is or how it relates to others. It needs a useful numerical description that can be transformed during computation.
Token Embedding converts the identifier of a token — a vocabulary piece of text — into a vector, a list of numbers. Typically the number identifies a row in a learned table. If the token's number is 17, we read the corresponding seventeenth row, rather than treat 17 as a measure of meaning.
Table values are adjusted during learning. The same token initially gets the same description regardless of the sentence. Later layers can then transform it using context: the Polish word “zamek” in “zamek w drzwiach” (“a lock in a door”) and “zamek na wzgórzu” (“a castle on a hill”) needs different surroundings.
An embedding is therefore neither a ready word definition nor proof of understanding meaning. It is an input description used by later mechanisms, such as Self-attention.
Mechanism source: the Keras Embedding layer.
Mechanism and details
From ID to vector
Imagine a portion of a table with three coordinates:
| ID | Vector |
|---|---|
| 6 | [0.2; −0.1; 0.5] |
| 9 | [−0.4; 0.3; 0.8] |
For input [6, 9, 6], we retrieve the first, second, and then the first of the displayed rows. These are invented values, and semicolons separate coordinates. A repeated ID yields the same vector from the same table, even if it occurs elsewhere in the sentence.
Token Embedding does not yet contain the result of analyzing neighboring tokens. Self-attention can later give two occurrences different context-dependent representations. Positional Encoding separately provides information about position in the sequence.
Lab: from token to training
Below, you can follow table lookup, rename IDs, change a coordinate, and perform a training step. All four sections use the same table. This original example has six tokens and six coordinates per token; words are treated as indivisible units. A real tokenizer may split words into smaller pieces.
Where vector values come from
In the original Transformer, embeddings were trained together with the model; prior Word2Vec training was not required. This is described in Attention Is All You Need, §3.4.
Table with dimensions stores parameters: is the vocabulary's token count, and is the length of each vector. Looking up three IDs produces a matrix. Renaming IDs while rearranging the corresponding rows preserves the lookup result. Changing a row's contents, however, changes the vector retrieved for every occurrence of that token. The table's shape and trainable nature are also described in the PyTorch Embedding documentation.
In the training section, the input is “kamienny,” and the correct next token is “zamek.” The model has only two possible responses: “zamek” and “klucz.” Two fixed weight vectors transform the input embedding into numerical scores called logits. Softmax converts them into probabilities. Cross-entropy for this sample is ; it is a special case of the loss described in the PyTorch documentation.
The training button calculates the gradient of this loss with respect to the six coordinates of the input vector and subtracts the gradient multiplied by 0.4 from them. With the initial settings, the first step raises the target probability from about 47.0% to 61.0% and reduces the loss from about 0.755 to 0.495. The “kamienny” row changes because it was the input; “zamek” was the target label. These numbers come from the example's calculations.
In this experiment, output weights are frozen and separate from the input table, with no momentum [Polski] or weight decay. This exposes only the update of one accessed row. In a full model, the training method and weight sharing may also change other parameters. Improvement on one sample demonstrates optimization, rather than the quality of a language model.
Where context enters
“Oto kamienny zamek” and “oto metalowy zamek” use the same token, “zamek,” at the same position in the lab. Table lookup yields an identical vector. Adding the same Positional Encoding also preserves equality between the two inputs. Only the Self-attention calculation accounts for the difference between the preceding tokens.
The demonstration calculates one attention head with identity projections for Q, K, and V, and sinusoidal Positional Encoding according to Attention Is All You Need, §3.2 and §3.5. You can inspect all attention weights and make the sentences identical so that the results match again. This models the effect of context with fixed weights; it is not a trained classifier of the two meanings of “zamek.”
The article on attention [Polski] discusses later vector transformations in more detail. Its term “updated embeddings” refers to representations after computation, which must be distinguished from table lookup itself.
Sources
- Keras: Embedding layer — converting IDs into vectors.
- PyTorch: Embedding — the parameter table and output shape.
- PyTorch: CrossEntropyLoss — loss for a target expressed as a class ID.
- Vaswani et al., 2017, Attention Is All You Need — embeddings, Self-attention, and Positional Encoding.