Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning
Argus is a persistent agentic runtime for long-horizon work in which model weights remain fixed while runtime state and control policy evolve. Manager, Planner, Engineer, and Reviewer roles execute bounded missions over durable project state, separating user intent from operational objectives, constraints, and verification criteria. The abstract reports roughly 78% on SWE-Bench Pro versus 59% for Direct Copilot at 1.41x aggregate tokens, plus 21% fewer solve-input tokens and 15% less active workflow time in mature waves. It also reports results on AARRI-Bench, mathematical data synthesis, GPU kernels, language-model training, and multi-day research campaigns.
Why it's worth reading
As agent systems move beyond single-turn generation, persistence, recovery, and verified accumulation become central bottlenecks; Argus offers a concrete runtime design and reports measurable results across software, research, and engineering tasks.