Mastermind treats strategy, rather than the complete action trajectory, as the learning unit for repository-scale vulnerability reproduction. Its dual-loop design combines a trainable planner, optimized with supervised fine-tuning and milestone-based GRPO, with a task-local experience loop that records reusable strategies. On CyberGym, the system uses 260 training tasks and 200 held-out evaluation tasks. With GPT-5.5 frozen as the executor, it reaches an 84.5% pass rate, compared with 60.0% for open-book PoC context, 63.0% for Best-of-8 sampling, and 77.0% for iterative improvement. The same planner also improves GPT-5.4 mini and GLM 5.1.
No heat snapshots are available in the last 24 hours.