AREX introduces recursively self-improving agents for deep research. An inner loop gathers evidence and builds a provisional answer, while an outer loop audits the answer constraint by constraint, identifies unresolved claims, and launches targeted follow-up research. The system also learns an autonomous context-update tool that compresses long interaction histories into an improvement state containing verified evidence and unresolved constraints, without using an external model. The authors instantiate both a dense 4B model and a 122B-A10B mixture-of-experts model, reporting substantial gains over comparable-scale baselines across BrowseComp, WideSearch, DeepSearchQA, Humanity's Last Exam, and other benchmarks.
No heat snapshots are available in the last 24 hours.