This paper proposes verification as a new scaling axis alongside pre-training, post-training, and test-time compute. Its LLM-as-a-Verifier framework estimates continuous scores from the expectation of scoring-token logits rather than producing discrete judge scores. Verification can scale through finer score granularity, repeated evaluation, and decomposed criteria. The paper reports state-of-the-art results on Terminal-Bench V2 (86.5%), SWE-Bench Verified (78.2%), RoboRewardBench (87.4%), and MedAgentBench (73.3%). It also describes extensions for Claude Code, agent-progress estimation, and dense reinforcement-learning feedback.
No heat snapshots are available in the last 24 hours.