CodeRescue treats a coding agent’s post-failure choice as a recovery-routing problem over heterogeneous actions: continue with a cheap model, spend more cheap compute, or escalate to a stronger model. The authors train a supervised router from execution rollouts and add a Conformal Risk Control (CRC) layer that chooses a deployment-time cost penalty without retraining. On held-out failures from five coding benchmarks, cheap recovery and escalation show complementary success patterns. In the main GPT-5.4-nano/GPT-5.4 experiment, one CRC-calibrated frontier point reportedly exceeds always-escalate solve rate while using 35% of its mean recovery cost.
No heat snapshots are available in the last 24 hours.