Back to archive
#ai#papers#aigen#llm#inference

Quantization errors partly cancel out across model layers

You want to run a language model on your own graphics card, but its weights take up too much memory. One solution is post-training quantization: you replace numbers stored in the finished model with values from a smaller set. You save space at the cost of numerical precision.

However, these numbers participate in the calculations for every subsequent part of the response. The model predicts the next token, a unit of text, using the earlier text and its weights. After quantization, it has approximations of those weights. An error can change the selected token, at which point the rest of the response is generated from a different history.

You might therefore expect deviations to accumulate as they pass through successive layers. Yet the compressed model often retains useful quality. Knowing that each rounding error is small does not explain what happens to its effects dozens of operations later.

Inside the model, errors have directions. A deviation produced by the current block can partly cancel a deviation inherited from earlier calculations. This relationship is studied in the preprint “Why Does Post-Training Quantization Work?”, in the field of model compression and LLM inference — running models after training. The authors also examine what happens when such compensation is deliberately removed.

A block processes input already changed by previous layers

A Transformer stores an intermediate computational result as a vector, an ordered list of numbers. The vector describes the text processed so far and passes through successive blocks. Through a residual connection, a block adds its own result to the input it receives. It therefore passes on a sum: input plus the calculated update.

Compare two runs for the same text: the model with its original weights and its quantized version. Their vectors already differ at the input to a selected block. The block also calculates two different updates because both its weights and its input differ. The update difference includes both effects; it cannot be equated solely with a new weight-rounding error.

After addition, the output difference is the sum of the input difference and the update difference. This is simply a consequence of the residual connection. If the two differences point in partly opposite directions, their sum can be smaller than the error inherited from the input. During normal operation, the model does not calculate these differences or compare itself with a more precise copy. Researchers perform that comparison by running both versions side by side.

The same error magnitude can produce a different sum

Take a teaching example with vectors containing two numbers each. These numbers are prepared for the example, not measured in the paper. The sign of each coordinate indicates the direction in which that result deviates from the reference value.

Stage of one blockOriginal weightsAfter quantizationDifference: quantized minus original
Input(10, 10)(13, 10)(3, 0)
Update calculated by the block(0, 2)(-1, 3)(-1, 1)
Output: input plus update(10, 12)(12, 13)(2, 1)

In the study, the input is the hidden state at the block boundary, and the update is the result added through the residual connection. The two columns correspond to runs of the same model before and after quantization. The input difference (3, 0) and the update difference (-1, 1) sum to (2, 1). The first coordinate moves closer to the reference result, although a new deviation appears in the second.

We calculate the deviation's magnitude as the sum of squared coordinates. Before the block, it is 3² + 0² = 9; afterward, it is 2² + 1² = 5. Here, a smaller value means a smaller difference between the runs. This is an internal error measure, with no direct conversion into the percentage of correct answers.

Now change only the direction of the update difference. Replace (-1, 1) with (0, √2), where √2 is the square root of two. Both vectors have the same sum of squares, 2, but the new one is perpendicular to the input difference. Addition gives (3, √2), with a sum of squares of 9 + 2 = 11. We removed the component that reduced the first deviation while preserving the magnitude of the added difference.

This last operation corresponds to the authors' intervention: they remove the component of the update difference parallel to the input error, then restore its original length. In the real model, vectors have many coordinates, and the intervention covers a series of blocks. At each step, the update is recalculated from the already modified input. Our single step omits those later calculations and changes in the reference vector's length.

Removing compensation increases error with the same weights

In the main experiment, the authors use an existing Qwen3-32B and round selected weights to a four-bit format. They do not fine-tune the model or select rounding on calibration data. Ordinary quantization lowers mean accuracy on six answer-selection tests by 0.43 percentage points compared with the original weights. The exact tasks and evaluation procedure are described in the appendix below.

A separate experiment examines why errors accumulate slowly. On web-text passages from C4, the authors apply the intervention described above while retaining the same trained model and quantized weights. The final relative hidden-state error becomes 2.94 times larger than with ordinary quantization. “Relative” means the length of the vector difference divided by the length of the reference vector. In both comparisons, the error reference is the unquantized model.

This is a diagnostic intervention: researchers change internal computations to test the role of error direction. The result supports the claim that partial cancellation of differences limits their later growth. It does not mean an almost threefold change in answer accuracy. The intervention itself also requires an additional reference run, so it is not proposed as a cheap inference technique.

The authors observe this compensation increasing during training by analyzing successive saved weights. They do not establish which training property causes it. Assuming every future model will automatically learn to cancel errors equally well would go beyond the paper's result.

The output vector and token selection are two different measurements

The final vector reaches the LM head, the layer calculating a numerical score for each vocabulary token. Every token has its own weight vector. Its score is produced by multiplying corresponding coordinates and summing the results. The effect of a hidden-state change therefore also depends on its direction relative to a token's weights.

In the models studied, scores for highly ranked tokens are more stable after quantization than scores for lower-ranked tokens. The authors explain this observation through vector geometry and compare theoretical approximations with measurements. The derivation, however, assumes a uniformly random direction of hidden-state rotation. It provides no guarantee for an arbitrary direction of real error.

Even small score changes can swap two closely matched candidates. Stable average quality and identical response text are separate requirements. In an application expecting a specific tool call, the latter matters too.

Complete generated responses must be tested before deployment

The study's main limitation comes from shared text history. In token-level analyses, both model versions receive the same preceding tokens. The experiment does not track how the first different choice changes a long, freely generated response. Likewise, the geometric analysis describes average behavior; it does not rule out rare large deviations.

Experiment details

Source and checkpoint. Yuxiang Chen, Michael Beyer, Jun Zhu, and Jianfei Chen, “Why Does Post-Training Quantization Work?”, arXiv:2609.11716v1, first submitted September 10, 2026. The paper is a preprint. The main comparison uses a published, trained Qwen3-32B checkpoint, a model with 32 billion parameters. There is no post-quantization training stage or selection of the best checkpoint based on quantization results. Comparisons with random weights and earlier training stages are separate diagnostic analyses.

Weight format. The reference is BF16 [Polski], a sixteen-bit floating-point format. The W4 variant uses NVFP4 and RTN, rounding to the nearest representable value. Quantization covers projection matrices in self-attention, the mechanism exchanging information between text positions, and in the MLP, the network processing each position's vector. Token input vectors, normalization layers, intermediate results, and the LM head remain at the original precision. NVFP4 combines four-bit values with an additional scale for each group of 16 elements and a scale for the entire matrix. The authors simulate rounding and reconstruct weights for computation. They do not measure the speedup or memory usage of an engine operating on packed four-bit weights.

Texts and aggregation. Standard token-level analysis covers 64 randomly selected sequences with 512 prediction positions per dataset: C4 (web text), WikiText-103 (Wikipedia), and GSM8K (math problems with reference solutions). In GSM8K, the model predicts the next token from supplied correct text, so the result does not measure independent problem solving. Statistics are averaged first over positions, then over inputs. Error bars in token-level analyses show the standard deviation between complete inputs, not uncertainty in the mean.

Intervention. The 2.94× result comes from Qwen3-32B on C4, with RTN NVFP4 and compensation removed in blocks 17–48; the comparison is ordinary W4, and the measure is final relative hidden-state error (Figure 3A, Section 3.3, Appendix E.7). The length of the update difference is preserved locally when it is modified. Later blocks receive a changed state, so their differences may already have a different length from those in the run without intervention. The full study also includes direction reversal for negative interaction and measurements of changes to token probability distributions. Building these runs requires reference, quantized, and additional intervention computations; the publication does not give a total GPU-hour cost.

Filters. The calculation of error components in Figure 2 excludes the first sequence position and trajectories in which the relative difference in state lengths exceeds 0.5 at any depth. The second filter removed 0.08% of Qwen3-32B positions on C4. The appendix also shows results without this filter. This selection should not be transferred to answer-accuracy tests, nor should analysis of typical trajectories be treated as a guarantee for extreme cases.

Answer tests. The mean 0.43 pp drop applies to Qwen3-32B, W4 versus BF16, without task examples in the prompt and without a chat template. The full test sets of ARC-Challenge, ARC-Easy, and MMLU were used, along with the validation sets of HellaSwag, WinoGrande, and TruthfulQA. They cover knowledge questions, text-ending selection, pronoun-reference resolution, and true-answer selection. ARC and HellaSwag evaluate candidate likelihood normalized by character count, MMLU evaluates answer letters, WinoGrande evaluates the shared sentence suffix conditioned on the candidate, and TruthfulQA uses MC1 accuracy. This evaluation is separate from hidden-state analysis. On C4, ordinary W4 changes the highest-scoring token at 10.7% of evaluated positions despite small average quality changes (Appendix C.6).

Scope of additional checks. The paper covers Qwen3-4B, 8B, 14B, 30B-A3B, and 32B, OLMo3-7B and 32B, OLMoE-1B-7B, Gemma3-4B, and Pythia-1.4B and 2.8B. Error-growth analysis compares Qwen3-32B with ten random initializations; the random reference for LM-head geometry uses three. Six OLMo3-7B Stage-1 checkpoints and Pythia training stages were also studied. Appendices include activation-only quantization, joint weight-and-activation quantization, and GPTQ and AWQ, methods that select quantization on calibration data. These are not the settings for the 2.94× result. The GPTQ and AWQ comparison used 64 C4 calibration inputs and eight disjoint evaluation inputs; RTN does not require calibration.

Yuxiang Chen, Michael Beyer, Jun Zhu, Jianfei Chen, Why Does Post-Training Quantization Work?, full text v1. Mechanism: Sections 3–4; settings: Appendix D.3; interventions: E.7; limitations: 5.4. Metadata and version history, DOI: 10.48550/arXiv.2609.11716. The source is available under CC BY 4.0. This is an English translation of a Polish discussion with an original numerical example.

My recommendation when choosing a compressed model version: measure success on the entire task using representative prompts, including long responses and tool calls. Comparing weights alone or the mean accuracy of short tests does not examine the effects of diverging generation histories. It is worth separately comparing the correctness of tool calls and final answers under an identical generation procedure.