LoRA
You have a large model and want to adapt it to the style of your documents. Training all its weights — the numbers controlling its computations — requires a lot of memory. Can you train only a smaller correction to what the model already does?
LoRA (Low-Rank Adaptation) retains the base weights and trains two smaller matrices, tables of numbers describing their change. The product of these tables gives a correction added to a selected model matrix.
Instead of changing every cell of the large table independently, the update passes through a narrow intermediate dimension. Its size, called rank r, limits the number of independent directions of the correction. The saving concerns trained parameters, rather than removing the base model from memory.
A small r provides a cheaper but less flexible way to adapt the model. LoRA is a variant of Fine-tuning; it does not guarantee quality without good data and settings. The example below lets you count the numbers being trained and distinguish that count from total memory requirements.
Mechanism and details
In the original parameterization, for a base matrix , we use:
and contain the trainable parameters. A small limits the rank of the update, and sets its scale. The restriction concerns the change in weights, not the rank of the entire base matrix. The mechanism is described by Hu et al., LoRA: Low-Rank Adaptation of Large Language Models, §4.1.
How many parameters do we actually save?
An original example: a 4096 × 4096 matrix has 16 777 216 parameters. For , the two LoRA matrices contain a total of parameters. That is 256 times fewer trainable parameters for this one matrix, assuming we train only its adapter.
The base weights still need to be stored. Activations and other computational elements also need memory, so this example does not imply 256 times less VRAM usage. In §4.2, the authors separately describe savings in optimizer state [Polski] and saved adapter size.
In a Transformer, LoRA can be applied to selected attention projections, but the method is not restricted to one pair of matrices. The final scope of training depends on the configuration: rank, selected layers, and any additional trainable modules. The PEFT documentation describes, among other things, r, target_modules, and modules_to_save.
An older post about the DeepSeek R1 training script [Polski] contains an example LoRA configuration. It is related material; the definition above is based on the original paper and documentation.