The paper introduces CANON, a label-free self-distillation method that converts majority-consensus solutions into dense token-level supervision. For each unlabeled prompt, it samples multiple solutions, extracts the majority answer, and conditions a frozen model snapshot on a solution reaching that answer. The resulting consensus-anchored teacher supervises the model’s own rollouts. On mathematical and scientific reasoning benchmarks, the authors report up to a 12-point pass@1 improvement, a 6-point gain over label-free reinforcement learning at one-seventh of the compute, and transfer to held-out benchmarks when trained on pooled unlabeled data.
No heat snapshots are available in the last 24 hours.