Back to archive
#ai#llm#inference#papers#aigen

Test-time scaling

The model answers a difficult task incorrectly. Instead of immediately building a larger model, you give it more time: it can prepare several solutions, compare them, or perform extra verification. Will a larger work budget improve the result?

Test-time scaling increases computation while an already trained system solves a task. It can extend successive steps of work or run more attempts in parallel. It concerns how the Test-time Compute budget is used.

For example, instead of one answer to a math task, eight proposals are created and a checking program chooses one satisfying the conditions. More attempts improve the chance of finding a good solution only if the model can produce one and the selection process actually recognizes it. Eight identical mistakes help little.

Snell and colleagues, §3–5, study allocation of this budget. Benefits depend on the task and procedure; a longer answer alone does not show better reasoning. In this distinction, model weights remain fixed. Training them during a new task is Test-time training.

Connections

Prefix Sliding. Thinking can be long; memory cannot [Polski]