The paper introduces SeGaBench, an executable benchmark with 100 synthetic and 20 source-backed C/C++ cases for testing whether LLMs can recover optimization-enabling semantics unavailable to conventional compiler analysis. Each case includes hidden semantics, an oracle artifact, validators, and a reproducible performance protocol. Across five models and five responses per case, the strongest model reportedly generated correct artifacts in 94.8% of responses; 83.3% reached at least 1.05x speedup, and 93.3% of cases had a performance success. Correct outputs often captured only part of the oracle’s potential.
No heat snapshots are available in the last 24 hours.