Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?

First seen · 7/2/2026, 12:00 PMLatest activity · 7/2/2026, 12:00 PM

This paper audits GSO, SWE-Perf, and SWE-fficiency by replaying official reference patches for 740 repository-level optimization tasks across four Google Cloud machine types. Only 39/102 GSO, 11/140 SWE-Perf, and 411/498 SWE-fficiency tasks satisfy their original validity rules in every cross-machine replay. Rankings also vary materially with scoring design: GSO and SWE-fficiency disagree on 9 of 28 pairwise comparisons, while SWE-fficiency’s ten most influential tasks receive 58.5%-82.8% of total score weight. The results show that aggregate leaderboards can conflate agent capability with runtime noise, task saturation, and benchmark-specific weighting.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/2, 12:00 PMnot independentRepresentative
    Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?