Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

Read original
HuggingFace Daily Papers·Fanzhe Meng·Aug 5, 2026, 8:00 PM

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

Papers72

CalibForge synthesizes terminal-agent training tasks and revises them using verified solver behavior. Its multi-solver strategy seeks disagreement across heterogeneous solvers, while contrastive calibration enforces a strong-solver-pass and weak-solver-fail relationship. The authors report 5,431 calibrated tasks, with trained models scoring 32.58% and 47.57% on Terminal-Bench 2.0. Reported gains over corresponding base models reach 24.71 percentage points on Terminal-Bench 2.0, 27.68 on SWE-bench Pro, and 30.04 on Doc2Repo. The supplied arXiv identifier and publication date are future-dated and therefore require verification.

Why it's worth reading

The work offers an actionable way to target solver-relative task difficulty, but its unusually future-dated arXiv metadata should be verified before relying on the reported results.

Tags

CalibForge终端智能体任务合成求解器校准Terminal-BenchSWE-bench Pro训练数据

Score breakdown

  • Novelty84
  • Impact78
  • Practicality82
  • Credibility48
  • Timeliness68