Order-consistency SFT
You show the model the same documents in two orders. One document's score changes from 0.70 to 0.60. If you then discard documents below a threshold, shuffling alone can change the system's answer. You want to limit this instability during training.
Order-consistency SFT is further training of a scoring model that encourages consistent scores for the same candidate across different orders. SFT means supervised fine-tuning: adapting a model using examples with specified results.
In our example, the mean of the two scores is 0.65. An additional penalty accounts for both scores' deviations from this mean and encourages them to move closer. Assessment against a sensible label is also needed: a model always returning 0.65 would be consistent but would not distinguish good documents from bad ones.
The authors, §3, combine these two training objectives. Shuffles are used during training and need not be repeated for every later query. The method limits order effects, but guarantees neither their complete removal nor correct scores outside the tested scope.