This paper audits GSO, SWE-Perf, and SWE-fficiency by replaying official reference patches for 740 repository-level optimization tasks across four Google Cloud machine types. Only 39/102 GSO, 11/140 SWE-Perf, and 411/498 SWE-fficiency tasks satisfy their original validity rules in every cross-machine replay. Rankings also vary materially with scoring design: GSO and SWE-fficiency disagree on 9 of 28 pairwise comparisons, while SWE-fficiency’s ten most influential tasks receive 58.5%-82.8% of total score weight. The results show that aggregate leaderboards can conflate agent capability with runtime noise, task saturation, and benchmark-specific weighting.
No heat snapshots are available in the last 24 hours.