DataComp for VLMs (DCVLM) introduces a controlled benchmark for evaluating data curation strategies for vision-language model training. It collects 160 datasets across image-caption pairs, interleaved multimodal documents, text-only data, and instruction-tuning data, totaling 6T multimodal tokens. The benchmark covers 1B–8B models, training budgets from 6.25B to 200B tokens, and up to 52 downstream benchmarks across nine domains. Experiments report that data mixing matters more than filtering, with instruction-heavy mixtures scaling better than caption-heavy ones. The DCVLM-Baseline reaches 63.6% on a 33-task core suite, 5.4 points above FineVision.
No heat snapshots are available in the last 24 hours.