CalibForge synthesizes terminal-agent training tasks and revises them using verified solver behavior. Its multi-solver strategy seeks disagreement across heterogeneous solvers, while contrastive calibration enforces a strong-solver-pass and weak-solver-fail relationship. The authors report 5,431 calibrated tasks, with trained models scoring 32.58% and 47.57% on Terminal-Bench 2.0. Reported gains over corresponding base models reach 24.71 percentage points on Terminal-Bench 2.0, 27.68 on SWE-bench Pro, and 30.04 on Doc2Repo. The supplied arXiv identifier and publication date are future-dated and therefore require verification.
No heat snapshots are available in the last 24 hours.