Opti-Agent-Bench evaluates LLM-based optimization agents across the full R&D pipeline: understanding business requirements, formulating mathematical models, selecting algorithms, implementing code, and producing solution reports. Unlike benchmarks built from pre-structured formulations, it introduces business-semantic authenticity and anti-template traps designed to resist pattern matching. Its modular evaluation uses cross-module consistency checks, while the ORAC bi-level validity framework addresses both task quality and scoring integrity. Across industrial-scale tasks involving integer programming, robust optimization, stochastic programming, and non-convex optimization, the benchmark exposes failures such as omitted constraints, model-code inconsistency, and divergence between reports and implementations.
No heat snapshots are available in the last 24 hours.