Back to archive
#ai#llm#glossary#aigen

Sequence Packing

You are teaching a model to write short customer responses: “Thank you”, “No”, and “Yes”. After division into pieces of text, or tokens, the examples illustratively occupy 3, 2, and 2 positions, including each text's end marker. If each is stored in a separate row of eight slots, only seven of the 24 slots contain data. The rest is padding, filler that aligns lengths but is not content to learn. How can the free slots be used?

Sequence Packing arranges several examples in one input sequence to reduce padding during training. In our example, all three fit into one eight-position row: seven slots with data, one empty. No new answers were created; their arrangement changed. The TRL 0.29.0 documentation, SFT Trainer, “Packing” describes this role.

Three colored paper strips closely arranged in one long tray, with thin dividers preserving their separation: a metaphor for Sequence Packing.

Eight slots are our own simplification, rather than TRL's default row size. An ordinary batch of examples may be padded only to its longest text; grouping similar lengths also reduces wasted space. Attention Is All You Need, §5.1, already describes this grouping. Empty slot count alone does not determine speedup: that also depends on how computations are executed.

A shared row need not mean a shared conversation

You want “No” to remain an independent example, rather than a continuation of “Thank you”. The end marker itself is only a token. Ordinary blocking of future access — Causal Masking — still lets later text read earlier text in the same row.

To preserve independence, computation must know example boundaries and block access across them. Error evaluation must be controlled separately: do not teach prediction of the first token of “No” from the previous text's ending. Choosing assessed positions is Loss masking; omitting a position's penalty does not block reading it.

Specifically, in TRL 0.29.0 code: DataCollatorForLanguageModeling and BFD packing configuration, example lengths reconstruct boundaries and beginnings are excluded from error evaluation. This path requires a compatible attention implementation — the mechanism for using other text positions; the code warns of information leakage in unsupported configurations. This is a requirement of the specific solution, rather than a property of all packing.

Check two separate things

Enable packing, then disable access isolation. The space saving remains, but B1 can use the earlier example A. Squares represent end tokens; A1, B1, and so on are illustrative text pieces. The demonstration shows allowed connections without simulating model responses. The “full rows” variant reveals no savings when every example already fills eight slots.

Strategies differ in how they handle long texts. In TRL 0.29.0, the pack_dataset function, bfd fits whole examples into free slots but truncates the excess of an overlong example; wrapped can split text in the middle. Both retained data and boundaries must therefore be checked. Packing organizes training inputs, including Fine-tuning — further model adaptation using examples. It does not increase the maximum text length the model supports.