AgentDebugX is an open-source framework for debugging LLM agents as a closed loop: Detect, Attribute, Recover, and Rerun. Its DeepDebug module combines global trajectory understanding, structure-guided investigation, and cross-examination for multi-turn root-cause diagnosis. On the Who and When benchmark, it reports 28.8% strict exact agent-and-step attribution accuracy with qwen3.5-9b, compared with 21.7% for the strongest single-pass baseline. On GAIA, one rerun repaired 13 of 73 failed tasks, versus 4–6 for three decoupled self-correction baselines, raising overall accuracy from 55.8% to 63.6%.
No heat snapshots are available in the last 24 hours.