arXivYiling Ma
IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
Papers77
Moving from an abstract research idea to faithful code often falters on omitted technical details. IdeaAMBIG benchmarks this codification readiness across 660 evidence-grounded gaps derived from GitHub issues and reproducibility artifacts. Across 13 evaluated language models, the strongest system achieved only a 9.6% defect recovery rate when locating omissions unaided. While models handle clarification well once a defect is explicitly marked, detecting what an author left unsaid remains their primary limitation.
Why it's worth reading
It exposes a critical bottleneck in autonomous research agents: while models can resolve defects once pointed out, they struggle fundamentally to detect unstated assumptions and missing specifications on their own.
Tags
Autonomous AgentsCode GenerationLLM BenchmarkReproducibilityAI for ScienceSoftware Engineering