CoRT addresses the coarse credit assignment in rubric-conditioned GRPO, where a response-level advantage is broadcast uniformly across all generated tokens. It replays the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt, then uses tokenwise log-likelihood contrasts as a proxy for rubric dependence. Bounded, response-normalized weights redistribute the signed GRPO advantage without changing the response-level reward or adding an auxiliary scorer. The paper reports improvements over matched response-level GRPO in the vast majority of comparisons, with an average gain of 4.4 percentage points, while remaining competitive with learned token-level credit baselines.
No heat snapshots are available in the last 24 hours.