This paper identifies two failure modes in rubric-based reinforcement learning for open-ended LLM tasks. Unexplored Criteria receive no signal because no rollout satisfies them, while Suppressed Criteria are satisfied by some rollouts but lose their learning signal after scalar reward aggregation produces non-positive aggregate advantages. The authors report that more than 57% of training samples exhibit suppression, with 1.8 suppressed criteria per sample on average. CriPO uses on-policy self-distillation to address both cases without train-inference mismatch, and outperforms rubric-based RL on medicine and science benchmarks with roughly twice fewer optimization steps.
No heat snapshots are available in the last 24 hours.