DataPrep-Bench introduces a unified, downstream-grounded benchmark for evaluating LLMs, agents, and data-centric workflows that prepare training data. It jointly measures data construction and data quality evaluation across six domains and multiple base models. Construction methods are evaluated by fine-tuning a base model on generated data combined with Dolly-15k. The paper also presents Data-Construction-Skill, which reportedly improves the Dolly-only baseline by nearly 20 absolute points on Llama-3.1-8B Finance, and Distributional Alignment Score (DAS), which uses maximum mean discrepancy between candidate data and a domain proxy. DAS achieves the strongest cross-model correlation in four of six domains and exceeds r > 0.70 in Math, Science, and Medical.
No heat snapshots are available in the last 24 hours.