Back to archive
#ai#llm#training#papers#aigen

Voting lies. Disagreement usually does not

On AIME 2026, the pseudo-label from Qwen3-1.7B voting is wrong for about 85% of tasks. It is still possible to learn from this, because among the responses that disagree with the cluster, about 79% are also wrong.

You want the model to improve at a difficult task now, while answering. There is no label. This is test-time training: learning on the same set on which the model is supposed to answer.

On-policy self-distillation (OPSD) learns token by token. A second pass of the same model receives the correct answer in the prompt and uses it to improve each step. Without a label, that answer is unavailable. GRPO receives one reward for the whole solution, so a wrong voting result corrupts the update once. If you feed that same wrong answer into distillation, the error enters every token.

Previous approaches and TTPO

A natural trick: sample K solutions and use the largest cluster as a pseudo-label. The TTPO paper (Wang et al., arXiv 2608.27448) calls this majority voting. Algorithm 1 uses plurality. 3/8 or 4/8 is not a majority.

TTRL uses this cluster as a single reward for the entire solution. If you replace the OPSD label with the same cluster and distill all solutions, the error enters every token.

TTPO splits the group into agreement with the cluster (P) and disagreement (N). It distills only P. It penalizes only N. Token weighting is part of the method, not a correction at the end.

LTTPO=OPSD(P)+λ GRPO(N)\mathcal{L}_{\mathrm{TTPO}} = \mathrm{OPSD}(\mathcal{P}) + \lambda\,\mathrm{GRPO}(\mathcal{N})

Example

In the experiment, the model samples K = 64 solutions. It uses eight trajectories for the update (K_train = 8, half from P and half from N). The eight cards below are an analogy for this second step, not the entire vote.

7, 7, 7, 12, 19, 42, 4, 31

The largest cluster is 7 (3/8). The truth is 42. 3/8 is not a majority. The cluster is wrong.

Agreement: the three sevens. Disagreement: 12, 19, 42, 4, 31. Four of the five disagreements are wrong. One is correct. A penalty for “not 7” is therefore usually appropriate (in the paper, about 79% of disagreements are wrong when the cluster is also wrong), but here it would reduce the weight of the only correct solution.

Mapping to the algorithm

CardsTTPO component
8 cardsK_train = 8, after sampling K = 64
7 with a count of 3the largest cluster
three sevensset P, OPSD loss
12, 19, 42, 4, 31set N, GRPO loss
42 among Nerroneous penalty: a correct solution in the negative set

The OPSD teacher gets the same 7 in its prompt that the three agreeing solutions have already produced. The gradient does not pull them toward a different answer. This amounts to distilling thinking mode into non-thinking mode.

Token weights: OPSD downweights positions where the student already agrees with the teacher (entropy and KL divergence, Soft-OR). GRPO retains the top 50% of tokens according to −log p times confidence and penalizes confident, low-probability errors. The bottom half, including locally correct steps, receives no penalty. In this branch, there is no positive advantage that would counterbalance such a penalty.

Below is the same set of eight. Green: P. Red: N.

On an easy task, the largest cluster is usually correct, and both moves help. On a difficult one, where TTT is needed, the split matters: distill only agreement, penalize disagreement, and allow for the possibility that N contains 42.

Results

Three separate settings.

OpenThoughts. TTPO does not use labels; OPSD does. Average across five benchmarks for Qwen3-1.7B: TTPO 40.1, OPSD 39.7.

Pure TTT, training on the test set, without annotations. Qwen3-1.7B on AIME 2026, HMMT 2026, and BRUMO 2025: 38.0% → 45.2%.

Evaluation without thinking mode after training on OpenThoughts, not after pure TTT. Gain over the base: +25.2 (1.7B), +30.6 (4B), +36.4 (8B).

The authors report the best checkpoint. TTPO: 100 steps, measured every 25.

An ablation on difficult AIME: when almost none of the sampled solutions reaches the truth, splitting by ground truth leaves P empty and GRPO advantage at zero. Splitting by the cluster always leaves P nonempty. This does not mean that an incorrect label teaches better than a true one. It means that on this benchmark, rare truth disrupts routing, while the cluster does not.

What this does not show

This is a recent preprint. The tasks are mathematics with an answer that can be automatically extracted and compared. Code without tests or open-ended text does not provide such a cluster.

When K is small or none of the sampled solutions is correct, plurality is noise. A penalty for disagreement sometimes hits the only correct solution.

The practical takeaway: do not inject the cluster's content into all solutions. Use the cluster only to split P and N. Distill agreement. Penalize disagreement. Allow for the fact that plurality is not a majority and that N may contain a correct answer.

Terms

src: https://arxiv.org/abs/2608.27448 TTPO: Test-Time Policy Optimization (Wang, Lu, Wang, Lv et al.)

AI [Polski]
Test-time scaling