The paper proposes Experiential Learning (EL), which turns an LLM evaluator from a scalar-reward judge into a coach. The coach converts assessments of on-policy responses into transferable experiential knowledge. A teacher model then conditions on this knowledge, while the policy internalizes it through on-policy context distillation. According to the abstract, EL outperforms rubric-based reinforcement learning across two policy families, whether feedback comes from the policy itself or a proprietary model. It also reports better generalization to held-out and unseen open-ended tasks and reduced reward hacking, suggesting that textual feedback may preserve distinctions that scalar rewards discard.
No heat snapshots are available in the last 24 hours.