The paper introduces Evolution Fine-Tuning (EFT), a mid-training method that converts evolutionary-search trajectories into supervision so language models can learn how to mutate, backtrack, and iteratively improve solutions. The authors build Finch Collection, containing 156K trajectories across 10 domains and 371 optimization tasks, and fine-tune open models ranging from 2B to 9B parameters. On 22 held-out tasks, EFT models reportedly outperform their base models by 10.22% on average. Combined with test-time reinforcement learning, EFT matches state-of-the-art results on two circle-packing tasks and improves over its base model on the Erdős minimum-overlap problem.
No heat snapshots are available in the last 24 hours.