The paper introduces AdaMAST, a method that induces compact, evidence-grounded failure taxonomies directly from an agent system’s execution traces, without hand-authored codes or human annotation. Codes are organized across system-level, role-specific, and domain-specific axes and reused as a feedback interface for agent-system search, runtime monitoring, and trajectory selection. The abstract reports improvements across five benchmarks, including SWE-agent resolution increasing from 60% to 70%, Claude Code from 64.0% to 70.7%, and Terminal-Bench 2.0 best-of-5 accuracy improving by 8–15 points over Pass@1.
No heat snapshots are available in the last 24 hours.