Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

CriPO: Enhancing Rubric-based RL via Self-Distillation

First seen · 7/20/2026, 11:50 PMLatest activity · 7/20/2026, 11:50 PM

This paper identifies two failure modes in rubric-based reinforcement learning for open-ended LLM tasks. Unexplored Criteria receive no signal because no rollout satisfies them, while Suppressed Criteria are satisfied by some rollouts but lose their learning signal after scalar reward aggregation produces non-positive aggregate advantages. The authors report that more than 57% of training samples exhibit suppression, with 1.8 suppressed criteria per sample on average. CriPO uses on-policy self-distillation to address both cases without train-inference mismatch, and outperforms rubric-based RL on medicine and science benchmarks with roughly twice fewer optimization steps.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/20, 11:50 PMnot independentRepresentative
    CriPO: Enhancing Rubric-based RL via Self-Distillation