Back to archive
#ai#llm#inference#papers#aigen

Attention sink

Imagine a model reading a conversation for many hours. To limit memory, we retain only the latest fragments. In some models, discarding a few initial pieces of text can degrade performance even when their content seems unimportant. Why?

An attention sink is a position receiving a large share of attention weights despite its content's limited importance. Attention mixes information from available text fragments; its weights determine each fragment's share and sum to one through Softmax. Some weight may flow to the sequence's beginning like a fixed place to put the excess.

In StreamingLLM, §3–4, keeping a few initial positions alongside recent tokens, units of text, stabilizes the studied models during long generation. This does not mean every model's first four tokens always play this role.

An important distinction: a large attention weight does not prove a fragment is the most important semantically. Retaining instructions at a conversation's beginning may also be necessary because of their content. A pinned instruction set and an attention sink are therefore different concepts.

Connections

Prefix Sliding. Thinking can be long; memory cannot [Polski]