EdgeBench studies how deployed agents improve through interaction with real-world environments. Across about 38,000 hours of interaction and 134 ultra-long-horizon tasks, the authors report that aggregate learning performance follows a log-sigmoid scaling law with R² = 0.998. They also observe that learning speed across model generations roughly doubles every three months. Tasks span scientific discovery, software engineering, optimization, professional knowledge work, formal mathematics, and games, with each supporting at least 12 hours of continuous operation. The release includes 51 tasks and the full evaluation framework.
No heat snapshots are available in the last 24 hours.