Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

First seen · 7/27/2026, 12:00 PMLatest activity · 7/27/2026, 12:00 PM

DataPrep-Bench introduces a unified, downstream-grounded benchmark for evaluating LLMs, agents, and data-centric workflows that prepare training data. It jointly measures data construction and data quality evaluation across six domains and multiple base models. Construction methods are evaluated by fine-tuning a base model on generated data combined with Dolly-15k. The paper also presents Data-Construction-Skill, which reportedly improves the Dolly-only baseline by nearly 20 absolute points on Llama-3.1-8B Finance, and Distributional Alignment Score (DAS), which uses maximum mean discrepancy between candidate data and a domain proxy. DAS achieves the strongest cross-model correlation in four of six domains and exceeds r > 0.70 in Math, Science, and Medical.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/27, 12:00 PMnot independentRepresentative
    DataPrep-Bench: Benchmarking LLMs as Training Data Preparators