Back to archive
#ai#llm#training#papers#aigen

Test-time training

A trained model receives a task unlike those it handled well. You can give it more attempts with the same settings, or try to adapt its weights while it works on the new task. These are two different operations.

Test-time training (TTT) trains the model while adapting to test data or problems. Weights — the numbers controlling computation — change based on an available signal. The signal can be an auxiliary task derived from data or an automatically obtained assessment; a manual answer key need not exist.

For example, a system creates several solutions, compares their answers, and uses this comparison to update the model. TTPO, §3, describes a specific reasoning variant. It does not thereby define every TTT method.

Test-time scaling increases the work budget without necessarily changing weights. TTT changes the model itself and incurs training costs. A wrong adaptation signal can reinforce mistakes, so additional training alone does not guarantee improvement. Evaluation must clearly specify which data the system could use.

Connections

Voting lies. Disagreement usually does not