Back to archive
#ai#llm#glossary#aigen

Position-wise Feed-Forward Network

The model gathered information from the rest of the sentence for a particular fragment. It now needs to process that description: combine features and compute a new representation. Reading other positions again is not the only operation required.

Position-wise Feed-Forward Network (FFN) is the part of a block that transforms each position's description separately. The description is a vector, a list of numbers. All positions in a given block pass through the same learned function, but obtain different results because they have different inputs.

In a small example, a list of four numbers expands to twelve, is transformed, and contracts to four. A nonlinear rule operates between expansion and contraction: for example, it keeps positive numbers and replaces negative ones with zero, so [−2, 3] gives [0, 3]. This lets the whole transformation do more than one operation summing numbers multiplied by constant coefficients. We share the computational rule rather than create separate weights for each word.

The FFN itself does not mix positions, although its input may already contain context from Self-attention. “Position-wise” describes this separate work, rather than requiring all later models to have identical structures.

Mechanism and details

The original Transformer used two affine transformations with ReLU between them:

FFN⁡(x)=max⁡(0,xW1+b1)W2+b2.\operatorname{FFN}(x)=\max(0,xW_1+b_1)W_2+b_2.

Here, xx is the vector of one position, W1W_1 and W2W_2 are weight matrices, and b1b_1 and b2b_2 are bias vectors. ReLU replaces negative coordinates with zeros. Parameters were shared between positions but not between successive blocks. Attention Is All You Need, §3.3.

The same function, different inputs

In a small invented model, a vector may pass through dimensions 4 → 12 → 4. For a sequence of 20 tokens, we perform the same function 20 times on different inputs. We do not create 20 separate sets of weights. If two input vectors are identical, the same deterministic FFN returns identical results.

The example in Dive into Deep Learning, §11.7.2 shows precisely this sharing of the function and the change in the tensor's last dimension.

The FFN output in a Transformer block is combined with the input through a Residual connection. The final dimension therefore returns to the dimension needed for addition. “Position-wise” describes how positions are processed, not a requirement to use ReLU in every later model variant.