The paper introduces Recursive Synthetic Terminal Tasks (RST), a verified recursive synthesis framework for long-horizon terminal-agent data. Starting from verified seeds, RST extends reference solutions, realigns instructions and verifiers, and validates each task in a fresh sandbox before reusing accepted tasks. Across 15 rounds, it generates 37,484 tasks at roughly $0.05 per task. Median solution length grows from 67 to 374 lines and executed commands from 40 to 244, while DeepSeek-V4-Pro pass@4 falls from 90% to 2.5%. Rejection-sampled training data improves Qwen3.5 models by up to 10 points on several terminal-agent benchmarks.
No heat snapshots are available in the last 24 hours.