ACPO is a token-level credit-assignment method for outcome-supervised reinforcement learning on verifiable reasoning tasks. It replaces global entropy with a mode-local proxy, one minus the probability of the top token, to reduce distortion from long-tail vocabulary probabilities. The method adds mismatch routing and saturation correction so that updates emphasize uncertain decisions on positive-advantage trajectories while penalizing confident regions on failed trajectories. The abstract reports consistent gains over entropy-aware methods such as 80/20 and GTPO, as well as outcome-supervised RL baselines including DAPO and SAPO, on mathematical and coding benchmarks including AIME 2025 and HumanEval Pro.
No heat snapshots are available in the last 24 hours.