Flux-OPD addresses open-ended language-model training where verifiable rewards are scarce by using contexts as preference supervision that evolves with student performance. The paper decomposes the reverse-KL objective and argues that the student is distilled toward the geometric mean of context-conditioned teachers, while a conflict term captures disagreement among those teachers. Flux-OPD injects contextual differences into a context-free teacher anchor and scales them by the conflict signal. The abstract reports improvements over existing OPD paradigms on open-ended tasks, but gives no numerical results, benchmark names, or implementation details.
No heat snapshots are available in the last 24 hours.