The paper argues that mixed heterogeneous tasks create distinct entropy regimes under the same LLM policy, making global or token-level entropy control inadequate. It also identifies an entropy-dependent comparability problem in GRPO-style normalized advantages. GEPO estimates entropy at the prompt-group level from existing grouped samples, derives adaptive thresholds from historical entropy statistics, and applies asymmetric advantage shaping: attenuating positive advantages in low-entropy groups to reduce over-exploitation, while weakening negative advantages in high-entropy groups to preserve exploration. The authors report consistent gains over GRPO and recent entropy-control methods across two base models and 13 benchmarks covering math, physics, science, code generation, and instruction following.
No heat snapshots are available in the last 24 hours.