The same ranking yields different decisions after reordering
A system answering questions from documents often operates in four steps. A search engine finds candidates, a scorer assigns each one a number indicating its usefulness, a threshold rejects low scores, and a reader model prepares an answer from the remaining documents. The user sees the answer, not the intermediate ranking.
If the scorer receives the same documents in a different order and changes a few scores near the threshold, the reader may receive different source material. The query and candidate pool remain unchanged. Only their place in the prompt differs, yet the system makes a different decision.
Reordering candidates is called reranking. It is usually evaluated by the quality of the document order, but such a metric may barely respond to swapping two similarly relevant items. The threshold responds to their absolute scores and may pass through a different set.
A shared prompt makes the score depend on position and neighboring documents
A classic pointwise scorer evaluates one document in one forward pass. Batched pointwise scoring puts several documents in a shared prompt and reads a separate score for each. The instruction, query, and rubric are then processed once for the entire group, reducing inference cost.
However, a shared prompt changes the computation itself. A document's score depends on its content, position in the window, and the other documents in that window. The position effect corresponds to the model's sensitivity to the position of information in a long context [Polski]. Here, the influence of neighboring candidates is added to it.
In the system studied, the model did not generate scores as text. The probability distribution of four rating tokens was read and used to calculate a continuous score. Tokens were not sampled, and temperature did not participate in the calculation. The order of candidates in the prompt changed.
A threshold turns a small score change into a different document set
The six documents below are an example prepared for this explanation, not data from the study. Documents C and D have the same true relevance grade. The scorer receives the same objects twice; only their order in the prompt changes. The threshold is fixed at 0.50.
| Document | True grade | Score in order 1 | Score in order 2 | Decision 1 | Decision 2 |
|---|---|---|---|---|---|
| A | 3 | 0.83 | 0.81 | keep | keep |
| B | 3 | 0.80 | 0.78 | keep | keep |
| C | 2 | 0.52 | 0.49 | keep | reject |
| D | 2 | 0.48 | 0.51 | reject | keep |
| E | 1 | 0.29 | 0.30 | reject | reject |
| F | 0 | 0.08 | 0.07 | reject | reject |
nDCG@10 measures whether documents with high relevance grades appear high in the ranking, with more weight on the first positions. Swapping C and D does not lower this metric because both documents have grade 2. The threshold, however, retains A, B, and C once, and A, B, and D the other time.
The Jaccard coefficient divides the number of shared elements by the number of elements in the union. Here, the intersection contains A and B, and the union contains A, B, C, and D, so the result is 2/4 = 0.50. The six rows correspond to candidates, each score column to one permutation, the 0.50 threshold to a frozen cutoff, and the two selected sets to retained sets passed to the next component.
Shuffling candidates reveals an error invisible to nDCG@10
The authors repeated scoring of the same candidates after random shuffles. For each collection, they tuned the threshold to F1 beforehand and froze it before the actual measurement. The reader models and the mechanism for selecting the chosen–rejected pair also stayed fixed. This meant successive outputs differed because of the scorer, rather than retuning a downstream component.
In the main passage-reranking experiment on 18 collections, Qwen3-4B scored windows of 20 documents across 10 permutations. Results were averaged over three seeds, and each variant's checkpoint was selected on a held-out split by ranking quality. The five trained variants differed by at most 0.010 nDCG@10 points. Retained-set overlap was 0.656 for single-order SFT and 0.835 for OC-SFT. Higher overlap means the threshold more often passed through the same documents after the input order changed.
This result applies to specific scorers using a shared prompt. It shows the separation between two measurements: nDCG@10 checks order, while retained-set overlap checks the threshold's decision. After OC-SFT, the reader answer in multi-document QA and the training pair in response ranking changed less often; the full settings and values are in the technical appendix.
OC-SFT penalizes score divergence between two orders
Ordinary single-order SFT fits the student model to the teacher model's ratings for one order. It contains no condition linking a document's score to its score after the window is shuffled.
Order-consistency SFT (OC-SFT) gives the student two shuffles of the same window. For each candidate, it calculates the mean of the two scores, and consistency loss penalizes the deviation of both values from that mean. The first view retains the standard loss against the teacher model's rating, so training still requires a correct relevance grade.
In the example, document C received 0.52 and 0.49. The mean is 0.505, so consistency loss moves both scores toward it. Ordinary SFT sees one of these values; OC-SFT gets both and directly reduces their divergence. It performs the same calculation separately for each candidate in the window.
Two views require roughly twice as many forward passes in a training step as one view. After training, the model still handles one order per request. The inference-time alternative, batched self-consistency, repeats scoring for many permutations while handling each request.
The method reduces the error but does not guarantee invariance
OC-SFT reduces dependence on order but does not eliminate it. Averaging several permutations during inference still improved the repeatability of models after training. The method moves part of the cost into training but does not provide a mathematical guarantee of identical scores.
The study covers English data, one answer skeleton, scorers without chain-of-thought, and windows containing at most 20 candidates. It does not test reasoning-intensive retrieval or longer prompt configurations. The threshold, reader, and preference model were frozen; a downstream component trained jointly with the scorer might learn to compensate for some variability. Two legal collections are proprietary.
The consistency weight and checkpoint were selected separately for each task on a held-out split solely by ranking quality. The results therefore do not come from one checkpoint and one weight applied to every setting without tuning.
The source is the arXiv v1 preprint Equal Ranking Quality, Different Decisions: Training Order-Consistent LLM Scorers by Markus Frohmann, Mahdiyar Alavi, Elizabeth Lingg, and Navid Rekabsaz, first published on August 27, 2026. Fields: computation and language, information retrieval, and machine learning; arXiv:2608.26762, DOI: 10.48550/arXiv.2608.26762.
Experiment details
- Passage reranking. Training used MS MARCO: about 30 thousand queries and 3 million document ratings. Evaluation covered DL19–DL23; eleven BEIR collections: NFCorpus, FiQA, Touche-2020, ArguAna, Climate-FEVER, TREC-COVID, DBPedia, SciFact, Signal-1M, TREC-NEWS, and Robust04; and proprietary Legal-A and Legal-B. Public candidate lists came from BM25, while the legal datasets had prepared pools.
- Multi-document QA. Training used HotpotQA, with zero-shot transfer to 2WikiMultiHopQA and MuSiQue. Each question had 10 passages, 2 of which supported the answer.
- Response ranking. Training used about 30.7 thousand UltraFeedback prompts. Evaluation covered RewardBench-2, Nectar, PPE-MATH, PPE MMLU-Pro, and RM-Bench.
- Scoring and measurement cost. In reranking, 100 candidates were divided into five windows of 20. One permutation required five batched forward passes; 10 permutations required 50. Ordinary inference with one order performed five such calls. At width 20, batching was about 45% faster per query than the pointwise variant.
- Training. The main model is Qwen3-4B. The authors used LoRA with rank 16, alpha 32, and dropout 0.05. OC-SFT used two views. Target consistency weights were 5 for reranking, 3 for QA, and 1 for response ranking; each weight and checkpoint was selected on a held-out split by ranking quality. Trained results are averaged over three seeds.
- Broader model test. The comparison covered 11 dense base models from 1.7B to 32B in the Qwen3, Gemma 4, and Granite 4.1 families, and the sparse Mixture-of-Experts Gemma-4 26B-A4B. OC-SFT had lower mean τ-PSI than order-averaged distillation for all 12 models.
- Study cost. Training the student for the main result from published silver labels took about 90 GPU-hours on one node with eight A100-40GBs. Full public evaluation requires 60–90 hours on one L40S-48GB; the entire study, including unreported trials, used about 100 thousand GPU-hours.
| Task | Single-order SFT | OC-SFT |
|---|---|---|
| passage reranking, 18 collections, B = 20 | nDCG@10 0.449; τ-PSI 0.209; retained-set overlap 0.656 | nDCG@10 0.459; τ-PSI 0.083; retained-set overlap 0.835 |
| multi-document QA, 3 collections, B = 10 | nDCG@10 0.951; τ-PSI 0.159; answer flip 0.177 | nDCG@10 0.961; τ-PSI 0.096; answer flip 0.125 |
| response ranking, 5 collections, B = 4 | nDCG@1 0.684; τ-PSI 0.333; pair flip 0.869 | nDCG@1 0.701; τ-PSI 0.201; pair flip 0.661 |
τ-PSI measures ranking disagreement between pairs of permutations: 0 means identical rankings, 0.5 no correlation, and 1 a reversed ranking. For τ-PSI, answer flip, and pair flip, lower is better; for nDCG and retained-set overlap, higher is better.
In a system using a threshold, a reader, or training-pair selection, a test before deployment should shuffle the same candidate pool and check retained-set overlap, answer flip, or pair flip. If the downstream decision changes, similar nDCG@10 is not a sufficient criterion for choosing a scorer.