DecoupleMix frames VLM pretraining data construction as two coupled but separable optimization problems: inter-class allocation across capabilities and intra-class dataset composition. It uses single-variable iterative search for class ratios, then scores datasets by Quality and Difficulty and solves a constrained convex allocation problem with a diversity objective. The paper reports consistent gains over heuristic baselines, transfer of ratios from small proxy experiments to larger scales without retuning, and competitive performance after 80B additional multimodal continue-pretraining tokens, although the abstract gives no detailed benchmark values.
No heat snapshots are available in the last 24 hours.