ReflectRL proposes reusing failed trajectories from stronger expert models as “Golden Negative Trajectories.” Instead of imitating these flawed demonstrations, the framework prompts reflective reasoning over their errors and then transfers the resulting behavior back into direct reasoning through a Reflective-to-Direct Policy Transition. The authors call the underlying effect the “Reflection Advantage”: on difficult tasks, critiquing a flawed attempt may be easier than solving from scratch. The abstract reports consistent gains across nine benchmarks, four LLM backbones, and four on-policy training methods with minimal overhead, but provides no exact performance or compute figures.
No heat snapshots are available in the last 24 hours.