This paper studies coding-agent failure as a temporal process rather than only a final outcome. The authors collect 3,843 execution trajectories from seven frontier models using OpenHands, MiniSWE, and Terminus2 on Terminal-Bench, then manually annotate 1,794 complete and valid trajectories containing more than 63,000 execution steps. Across 14 findings, failures are predominantly associated with epistemic errors, often begin within the first few steps, and may remain concealed until recovery becomes impossible. The results suggest that reliability work should emphasize early validation and intervention, alongside final task-success metrics.
No heat snapshots are available in the last 24 hours.