NexForge presents a requirement-driven framework for synthesizing executable agent tasks and expert trajectories without relying on predefined tools, repositories, or skill graphs. It starts from real-world capability demands, builds scenarios and task profiles, compiles task directives, provisions files and runtime dependencies, and collects expert rollouts for supervised fine-tuning. The paper reports that 3.6K terminal tasks and 2K office tasks improve Qwen3.5-35B-A3B Base from 22.5% to 52.0% on Terminal-Bench 2.0 and from 813 to 1338 Elo on GDPval. Scaling to 43.2K terminal tasks reportedly reaches 58.4%, while Nex-N2 models achieve 75.3% on Terminal-Bench 2.1 and 1585 GDPval Elo.
No heat snapshots are available in the last 24 hours.