Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma

First seen · 7/10/2026, 12:00 PMLatest activity · 7/10/2026, 12:00 PM

This paper proposes Unbounded Positive Asymmetric Optimization (UP) for reinforcement learning training of large language models. It argues that importance-sampling objectives face an exploration-stability dilemma: unclipped updates can destabilize training, while conventional clipping limits useful exploration. UP uses a stop-gradient anchor and applies asymmetric treatment: positive-advantage samples receive unclipped gradients, while negative-advantage samples retain standard clipping safeguards. The abstract reports compatibility with token-level GRPO and DAPO, sequence-level GSPO, and experiments across dense, MoE, and vision-language models. Detailed benchmarks and implementation evidence require inspection of the full paper.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/10, 12:00 PMnot independentRepresentative
    UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma