Back to archive
#ai#llm#glossary#aigen

Elastic Weight Consolidation

A model routes customer questions to “delivery” and “returns”. Now it has to learn to recognize complaints, but the old messages may no longer be stored. So they cannot be replayed during training. All that remains of the old task is the model itself: millions of numbers set during earlier training, called weights.

You could train the model on complaints with no restrictions at all. Then the weights shift wherever that helps the new task, including the ones that recognizing delivery and returns depended on. You could also penalize every change equally, but then the model clings to its old settings so tightly that it learns the new task poorly. What we need is an answer to the question: which weights may be moved, and which are better left alone?

Elastic Weight Consolidation (EWC) adds a penalty to training on the new task for weights moving away from their values after the old task, and the penalty is larger the more important a given weight was for the old task. The authors compare this to a spring attached to each weight: stiff for important weights, loose for the rest (Kirkpatrick et al., Overcoming catastrophic forgetting in neural networks, §2).

A row of discs attached by springs to posts: the discs on thick, stiff springs stay in place, while those on thin, loose ones shift toward a red circle.

Two weights, one new task

Take just two weights. After training on delivery and returns, both have the value 1. The first is very important for the old departments (importance 10), the second almost irrelevant (importance 0.1). The new task requires only that the two weights add up to 4. Many settings satisfy this: (2, 2), but also (1, 3).

  • Without protection, training moves both weights equally, to (2, 2). The new task is mastered, but the important weight has drifted by 1 and the loss on the old task rises from 0 to 5.05.
  • A uniform penalty on both weights stops them halfway, at (1.67, 1.67). The old task suffers less, but the sum is 3.33 instead of 4.
  • EWC shifts almost all of the movement onto the unimportant weight: (1.02, 2.80). The sum is 3.82, and the loss on the old task is only 0.16.

Loss is a number that says how wrong the model is; the smaller, the better. In the experiment below you can change the penalty strength λ and the kind of protection. The result is computed exactly for this two-weight example; it is not a simulation of a real network.

How we know which weight is important

In a real network, nobody supplies the importance in advance. It is estimated once the old task is finished, while its data are still available: you check how strongly a tiny change in a given weight changes the model's answers on the old examples. The paper uses the diagonal of the Fisher information matrix for this, that is, one importance number per weight. After that the data can be deleted; what remains are the old weight values and their importances.

During the new training, the sum of two terms is minimized (equation 3 in the paper):

L(θ)=LB(θ)+∑iλ2Fi (θi−θA,i∗)2L(\theta) = L_B(\theta) + \sum_i \frac{\lambda}{2} F_i \,(\theta_i - \theta^*_{A,i})^2

LBL_B is the loss on the new task, θi\theta_i is the current value of weight number ii, θA,i∗\theta^*_{A,i} is its value after the old task, FiF_i is its importance, and λ\lambda says how much the old task counts relative to the new one. In our example F=(10, 0.1)F = (10,\ 0.1) and λ=1\lambda = 1.

Limits of the method

The importance is an approximation computed at a single point, separately for each weight. The authors themselves show that their estimate can be overconfident about which weights are unimportant, and they call this the main weakness of the method (§2.2 and §3). In their experiments with Atari games, an agent with EWC learned many games one after another, but did not match separate networks trained for each game. The results concern digit recognition and games, not large language models. Our two weights and customer messages are our own simplification.

  • Catastrophic Forgetting [Polski] is a substantial deterioration of earlier skills after new training; EWC tries to reduce it without access to the old data.
  • Experience Replay solves the same problem differently: it stores some of the earlier examples instead of weight importances.
  • Continual Learning sets the goal, that is, learning successive tasks while retaining the earlier ones; EWC is one of the methods.
  • Fine-tuning is further training of a ready-made model; EWC changes its objective by adding a penalty for moving important weights.

The text and illustration were prepared with AI assistance. The two weights, their importances and the demonstration are our own simplified educational example; the illustration is a metaphor, not a diagram of the model's architecture.