The paper introduces Recursive Harness Self-Improvement (RHI), which represents an agent loop as a prompt-level harness specification and refines it iteratively using pairwise feedback over its own revision history. On 30 synthetic machine-learning research tasks covering quantitative finance, robotics, and pharmacy, a few iterations reportedly improved the performance ceiling of low-reasoning-effort agents beyond the corresponding maximum-reasoning-effort setting, while reducing inference cost by up to 60%. The authors attribute the gains primarily to better task-specific context management and inter-agent information flow, rather than longer reasoning traces.
No heat snapshots are available in the last 24 hours.