ParVL proposes parallel scaling for multimodal LLMs by reusing shared Vision Transformer and LLM backbone parameters across multiple vision and language branches, each differentiated through branch-specific prefix parameters. This expands inference compute without duplicating the full backbone and makes the allocation between visual encoding and language decoding adjustable. The authors report end-to-end, full-parameter supervised fine-tuning on roughly 13 billion tokens and improvements over single-branch baselines trained with the same recipe. Their experiments also indicate that the best vision-language compute allocation depends on the downstream task.
No heat snapshots are available in the last 24 hours.