The paper introduces SCOPE, a benchmark for evaluating LLMs on autonomous experimental design across 19 research domains. It is built from 300 recent papers from venues including ICML, NeurIPS, and ICLR, and evaluates both high-level planning completeness, covering main, ablation, and analysis experiments, and low-level configuration quality, including datasets, baselines, and metrics. The reported results indicate that most LLMs cannot directly produce high-quality designs, low-level configuration is a shared bottleneck, and search mode does not improve design quality. The authors also propose OptED, an agentic workflow using stage isolation, tools, and rule-based constraints.
No heat snapshots are available in the last 24 hours.