This paper frames test-time scaling as budgeted inference over an autoregressive model’s implicit prefix tree. It distinguishes three regimes: sequential scaling along one trajectory, leaf-level sampling followed by terminal voting or verification, and prefix-level search over unfinished states. The authors argue that evaluation should treat the full inference pipeline as the evaluated system, report protocol-matched compute and uncertainty, and separate end-to-end results from candidate-bank diagnostics. The abstract also describes reproducibility requirements and a planned release of more than 2 billion complete reasoning traces with verifier and token-level signals.
No heat snapshots are available in the last 24 hours.