Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Opti-Agent-Bench: Benchmarking End-to-End Optimization R&D Agents on Real-World Business Problems

First seen · 7/12/2026, 09:47 PMLatest activity · 7/12/2026, 09:47 PM

Opti-Agent-Bench evaluates LLM-based optimization agents across the full R&D pipeline: understanding business requirements, formulating mathematical models, selecting algorithms, implementing code, and producing solution reports. Unlike benchmarks built from pre-structured formulations, it introduces business-semantic authenticity and anti-template traps designed to resist pattern matching. Its modular evaluation uses cross-module consistency checks, while the ORAC bi-level validity framework addresses both task quality and scoring integrity. Across industrial-scale tasks involving integer programming, robust optimization, stochastic programming, and non-convex optimization, the benchmark exposes failures such as omitted constraints, model-code inconsistency, and divergence between reports and implementations.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/12, 09:47 PMnot independentRepresentative
    Opti-Agent-Bench: Benchmarking End-to-End Optimization R&D Agents on Real-World Business Problems