This paper studies failure localization in LLM-based multi-agent systems, where long-horizon, tightly coupled interactions make it difficult to determine both which agent caused a failure and where the trajectory first became irreversibly misdirected. It introduces AgentLocate, combining an LLM-based judge with independent evaluators, confidence-aware aggregation, and lightweight fine-tuning of the judge. The abstract reports consistent improvements over existing localization methods on two complementary benchmarks spanning tasks, agent configurations, and trajectory lengths, while maintaining token and runtime efficiency. However, the abstract does not provide the benchmark names, numerical gains, or ablation results.
No heat snapshots are available in the last 24 hours.