Harness-R1 treats runtime-harness editing as a learnable capability rather than a fixed prompt-engineering procedure. A dedicated 9B harness engineer turns batches of target-agent failures into validated executable patches, then receives rewards from fresh reruns of the frozen target agent on the same batch. The engineer is initialized with supervised fine-tuning and optimized online with group-relative policy optimization. On WebShop, ALFWorld, and DBBench, vanilla Qwen3.5-9B improves from 44.3% to 53.6%. After target-agent fine-tuning, a target-specific engineer adds another 5.0 percentage points, reaching 64.2% on average.
No heat snapshots are available in the last 24 hours.