Test-time Compute
Test-time Compute
The model needs to solve a difficult problem. You can request one short attempt or devote more work to proposals and checking them. Describing this cost requires a computation budget after training ends.
Test-time Compute is computation performed while solving new tasks. It includes the work needed to obtain an answer; discussions of increasing this budget concern, for example, extra reasoning steps, multiple attempts, or verification.
With eight proposed solutions, you also pay to produce and select among them. Longer text is not the only way to increase the budget, and extra attempts will not help if we cannot recognize a good one.
Test-time scaling studies how increasing this budget affects results. Snell and colleagues, §3–5, show dependence on the task and work allocation. There is no universal promise of fourfold improvement. Training weights during a new task is the separate concept of Test-time training.