TRIAGE addresses a weakness in GRPO for agentic reinforcement learning: assigning the final verifier outcome uniformly to every action token. A structured judge labels trajectory segments as decisive progress, useful exploration, no-progress infrastructure, or regression, then applies fixed bounded process rewards conditioned on those roles. The authors describe this as the optimal segment-level correction available from role labels alone, linking it to lower-variance policy gradients. On ALFWorld, Search-QA, and WebShop, TRIAGE reportedly improves success rates over GRPO for two policy models and beats scalar judge rewards and a shared-backbone value baseline. Completed ALFWorld and WebShop rollouts use 10.4% and 14.8% fewer environment-facing turns than GRPO.
No heat snapshots are available in the last 24 hours.