LongHorizon-Harness reframes long-horizon agent execution as explicit task-state management outside the model’s growing context. Its Manage-Execute-Audit loop uses a manager to select subtasks, a fresh-context executor to act, and a read-only auditor to verify environmental outcomes before state updates. The supplied abstract reports that Qwen 3.7-Plus improves from 51.8% to 80.7% on WeaveBench, 69.7% to 77.2% on Terminal-Bench 2.1, and 2.8% to 8.3% on OSWorld 2.0. However, the metadata is future-dated and truncated, so the paper and results require independent verification.
No heat snapshots are available in the last 24 hours.