Echoverse compiles specifications into stateful applications whose tasks are graded against the applications’ own databases. Its co-evolution loop uses every graded rollout both to repair environments, tasks, and verifiers, and to train the model. Across fourteen evaluation splits, a 9B model trained on twelve environments improves from 36.5% to 67.1%, coming within fourteen points of a much larger frontier model. The authors identify behavioral depth, failure-targeted task design, and environment improvement as key contributors, and release four environments with applications, seed data, and grounded graders.
No heat snapshots are available in the last 24 hours.