The paper proposes LLM-as-a-Tutor for reinforcement learning with non-verifiable rewards. One model acts both as an examiner, pairwise-comparing policy rollouts to identify prompts that no longer produce meaningful quality variation, and as a generator that appends atomic constraints to increase difficulty. The append-only design aims to keep prompt difficulty aligned with policy capability without an external schedule. According to the abstract, the method outperforms policy-unaware baselines and prior policy-adaptive approaches on three complex instruction-following benchmarks. The abstract does not provide benchmark names, model configurations, or quantitative gains.
No heat snapshots are available in the last 24 hours.