Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

When Does Muon Help Agentic Reinforcement Learning?

First seen · 7/20/2026, 12:00 PMLatest activity · 7/20/2026, 12:00 PM

This paper studies vanilla Muon versus AdamW for sparse-reward agentic RL on ALFWorld with Qwen2.5-0.5B-Instruct. Under Group-in-Group Policy Optimization (GiGPO), applying Muon only to hidden weight matrices increased final-window validation success from 0.290 to 0.546, while high-rate AdamW controls showed no post-update success. The benefit depended on the advantage estimator and learning rate: Muon improved GRPO at 3e-5, while GraphGPO with Muon at 1e-5 reached 0.901 success and improved normalized validation AUC. The evidence remains exploratory because comparisons use a single seed and one task.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/18, 01:49 AMnot independent
    When Does Muon Help Agentic Reinforcement Learning?
  2. AggregatorHuggingFace Daily Papers7/20, 12:00 PMnot independentRepresentative
    When Does Muon Help Agentic Reinforcement Learning?