ModelEquivBench evaluates LLM-generated optimization models through a seven-level semantic profile, E0–E6, covering construction, representation alignment, feasible-set relations, objective ordering, optimal-value equality, and optimizer-set equivalence. Each decision includes independently checkable evidence, such as replayable traces, explicit maps, exact-rational certificates, or counterexamples. Unsupported structures, incomplete searches, and resource limits are reported as UNKNOWN or N/A rather than guessed. On a frozen cohort of 173 base problems, the benchmark compares GPT-5.4, Claude Sonnet 4.6, and Qwen3.5-397B-A17B, revealing distinct failure stages that a single accuracy score obscures.
No heat snapshots are available in the last 24 hours.