This paper systematically studies how language, visual understanding, and visual generation interact during unified multimodal pretraining. Controlled experiments on synthetic and real-world data examine knowledge transfer, modality synergy, and architectural choices. The authors report that data complexity largely determines whether modalities cooperate or compete; shared attention and normalization with modality-specific feed-forward layers encourage synergy. Early joint unification outperforms late alignment or sequential training and avoids a reported “vision laziness” effect. The paper also proposes recipes reaching strong generative performance with 5% of the compute budget, followed by validation using multiple 13.5B MoE models trained on 2T tokens.
No heat snapshots are available in the last 24 hours.