Back to archive
#ai#llm#glossary#aigen

Continual Learning

A model routes customer messages to the “orders” and “returns” departments. It can recognize “Where is my parcel?” Now the company opens a complaints department. You train the model further on messages about damaged products. It gets better at routing complaints, but some delivery questions also start landing in the new department. Checking only complaints would give a false impression of progress.

You need to teach the model new things while preserving earlier abilities. This is the goal of Continual Learning: learning from data or tasks that arrive in sequence, while taking account of what the model has already learned. The name describes a problem and a family of approaches; it does not guarantee that forgetting will disappear. Kirkpatrick et al., Overcoming catastrophic forgetting in neural networks, §1 describe this need.

Circles and triangles remain embroidered on a single piece of fabric while a needle adds a new wavy pattern: a metaphor for developing skills while retaining earlier ones.

What changes in training and evaluation

Suppose the full archive of examples used for earlier training is no longer available. Before deleting it, we saved a small sample: 20 messages with their known correct departments. One possible method mixes this sample into the new complaint examples. Further training then also receives a signal to keep handling “Where is my parcel?” correctly. Reusing earlier examples is one of the strategies described by van de Ven and Tolias, Three scenarios for continual learning, §3.3.

The evaluation team keeps a separate test of old messages, which we do not use for training. After each update, we check both old and new categories. The numbers below are an invented example, not research results; each test contains 100 messages.

Model variantCorrectly classified old messagesCorrectly classified complaints
Before the new training90/1000/100 — this category does not exist yet
After training only on complaints40/10092/100
After training with a sample of earlier messages89/10087/100

The last variant preserves the earlier ability better, although it recognizes complaints less well than the second one. The table shows what to compare; it does not promise this result from adding 20 examples.

Conditions change the difficulty

In our example, the model chooses among all the departments it has learned. A hint such as “this message belongs to the new category” would make the test easier. Whether such hints and task boundaries are available matters when comparing methods (van de Ven and Tolias, §2–2.3). A small sample may miss rare questions; a good average does not prove that every ability has been preserved.

  • Catastrophic Forgetting [Polski] describes a substantial deterioration in old abilities; its demonstration shows how new training changes shared settings.
  • Fine-tuning means training a model further on selected examples. After successive training rounds, we also check its earlier abilities.
  • Knowledge Distillation can pass guidance from an older model to the one being updated; this is a different strategy from keeping the original messages (§3.3 of van de Ven and Tolias's paper).

The text and illustration were prepared with AI assistance; the messages, sample size and table results are our own educational example. The illustration shows the goal of preserving earlier abilities, not the model's structure.