The paper introduces Long-Horizon-Terminal-Bench, a benchmark of 46 long-horizon terminal tasks across nine categories, including experiment reproduction, software engineering, multimodal analysis, games, and scientific computing. Tasks are decomposed into fine-grained subtasks with dense rewards and partial credit. Across 15 frontier models, a task run averages 9.9 million tokens, 231 episodes, and 85.3 minutes. The strongest tested model reaches only 15.2% pass@1 at a partial-reward threshold of 0.95 and 10.9% at the perfect-reward threshold of 1.0, while mean pass rates are 4.3% and 1.7%.
No heat snapshots are available in the last 24 hours.