Back to archive
#ai#llm#glossary#aigen

Experience Replay

A model has learned to route customer questions to “delivery” and “returns”. Now you are training it further to recognize complaints, such as “The product arrived damaged”. If its next exercises contain only complaints, it may become worse at handling the familiar “Where is my parcel?”. The complete old archive is unavailable. Could it at least revisit a few saved messages?

Experience Replay (ER) means storing some earlier examples and reusing them for training. In a simple version, we mix examples from memory into a batch of new data. This memory is called a replay buffer: a container for practice data, not a record of everything the model knows. This joint training is described by Chaudhry et al., Continual Learning with Tiny Episodic Memories, §2, Algorithm 1.

Saved indigo and ochre cards return from a small box to a row of new red cards: a metaphor for practice during further learning.

What goes into the next exercise

Before losing the archive, we saved four messages along with their correct departments. We now have four new complaints. Without replay, the model trains on four new messages. With replay, it also practices on four old ones: it receives another signal that “Where is my parcel?” belongs to delivery. In this small example we use the entire buffer; larger buffers usually supply a sampled batch.

Counting memory slots is not enough, however. Four saved delivery questions contain no returns example at all. Replacing two of them with “How do I send my purchase back?” and “I want to return a shirt” leaves the number of replayed examples unchanged, but covers both earlier departments. The experiment below counts examples in a data batch; it does not predict model accuracy.

Replay can help only with available data

The model gets an opportunity to practice earlier answers, not a guarantee that it will retain every skill. A small random sample can entirely miss a rare category; researchers observed this problem with very small buffers (§4.4 of the same paper). Our messages are our own example, not a result of those experiments. The effect of training must be checked on separate old and new messages that were not used for training. The buffer also consumes memory and requires retaining selected data.

In games, in Reinforcement Learning — learning actions from rewards — the stored record includes a situation, an action, a reward and the next situation. Sampling such records reduces dependence on events that occur immediately one after another (Mnih et al., Human-level control through deep reinforcement learning, Methods, “Training algorithm for deep Q-networks”). These are different data from a message with its correct department; Atari results should not be transferred directly to LLMs.

  • Continual Learning sets the goal of further learning while retaining earlier abilities; ER is one way to pursue it.
  • Catastrophic Forgetting [Polski] describes a substantial deterioration of earlier skills, which replay may reduce.
  • Fine-tuning is further training of a model; replay changes the selection of its training data. Simply pasting old messages into a conversation is not this kind of model update.

The text and illustration were prepared with AI assistance. The messages and demonstration are our own simplified educational example; the illustration is a metaphor, not a diagram of the model's architecture.