This paper studies vanilla Muon versus AdamW for sparse-reward agentic RL on ALFWorld with Qwen2.5-0.5B-Instruct. Under Group-in-Group Policy Optimization (GiGPO), applying Muon only to hidden weight matrices increased final-window validation success from 0.290 to 0.546, while high-rate AdamW controls showed no post-update success. The benefit depended on the advantage estimator and learning rate: Muon improved GRPO at 3e-5, while GraphGPO with Muon at 1e-5 reached 0.901 success and improved normalized validation AUC. The evidence remains exploratory because comparisons use a single seed and one task.
No heat snapshots are available in the last 24 hours.