Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

First seen · 7/30/2026, 12:00 PMLatest activity · 7/30/2026, 12:00 PM

CoRT addresses the coarse credit assignment in rubric-conditioned GRPO, where a response-level advantage is broadcast uniformly across all generated tokens. It replays the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt, then uses tokenwise log-likelihood contrasts as a proxy for rubric dependence. Bounded, response-normalized weights redistribute the signed GRPO advantage without changing the response-level reward or adding an auxiliary scorer. The paper reports improvements over matched response-level GRPO in the vast majority of comparisons, with an average gain of 4.4 percentage points, while remaining competitive with learned token-level credit baselines.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/30, 12:00 PMnot independentRepresentative
    CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization