Post-training quantization
You already have a trained model, but it occupies too much memory to run on your chosen computer. Much of that space stores the numbers controlling computation. Can they be stored less precisely to shrink the model?
Post-training quantization (PTQ) reduces numerical precision after the main training is complete. In the weight variant, it replaces learned values with numbers from a smaller set that can be stored using fewer bits.
In a simplified format with step 0.1, weight 0.73 becomes 0.7. We save space but introduce error −0.03. More elaborate methods use data samples to choose rounding or scales that degrade results less. GPTQ, §3, describes one such method.
Fewer bits do not automatically mean proportionally less total memory usage or faster execution. An engine supporting the format is needed; scales and other elements also have a cost. PTQ is not arbitrary retraining of the model. The article on propagating differences between layers discusses how errors affect the answer.