The paper introduces STRACE, a framework for improving reflection-based optimization of long-horizon agents. It filters redundant failures at the batch level by mining failure patterns, then performs causal localization within each selected trajectory using a textual dependency graph. This removes non-causal steps and identifies the module responsible for the failure. On the formal verification benchmark VeruSAGE-Bench, STRACE improved the success rate of human-expert-designed agents from 42.5% to 58.5%, described by the authors as a 1.4x improvement. The authors also provide code on GitHub.
No heat snapshots are available in the last 24 hours.