Back to archive
#ai#llm#inference#papers#aigen

Sliding window

The model is creating a long response. It compares each next piece with earlier text, so computation and memory grow. What if we let it use only the latest part of the conversation?

Sliding Window Attention restricts a position's direct access to a neighboring window of tokens — pieces of text. In a model predicting a continuation, this usually includes the current position and a limited number of earlier ones. The window moves with successive positions.

If the window covers four positions, the fifth does not directly read the entire history, only the allowed ending. This reduces the number of connections in a single layer. Information from more distant positions can still arrive indirectly through successive layers; restricting the window is not exactly the same as removing all memory of earlier text.

Mistral 7B, §2, describes this mechanism and a rolling buffer. The cost of the saving is restricted access to distant content. A variant also retaining initial instructions has different connections from a sliding window alone; the article on Prefix Sliding discusses its context.

Connections

Prefix Sliding. Thinking can be long; memory cannot [Polski]