Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning

First seen · 7/20/2026, 12:00 PMLatest activity · 7/20/2026, 12:00 PM

Agon proposes competitive cross-model reinforcement learning in which two models solve the same problem while alternately drafting, reading, and competing against each other. Each model is rewarded for outperforming a rival that has seen its reasoning, creating implicit pressure to improve the trace without process labels or a learned reward model. The abstract reports that, on the hard split of DeepMath with Qwen3, Agon doubled GRPO pass@1, delivering roughly eight times the gain of an untrained Mixture-of-Agents baseline. The ordering reportedly replicated on competitive-programming code and across Qwen3.5 and Gemma 4.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/20, 12:00 PMnot independentRepresentative
    Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning