Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness

First seen · 7/30/2026, 07:17 PMLatest activity · 7/30/2026, 07:17 PM

DataClawEval introduces a benchmark for end-to-end data engineering agents using 100 production-grade tasks written by enterprise data engineers. The tasks span PySpark, MySQL, HiveSQL, PrestoSQL/Trino, and FlinkSQL, and are executed in isolated, case-specific sandboxes with deterministic rule-based grading rather than LLM judges. Across 16 frontier agents, the best overall score reaches only 74.9. No model dominates across all engines: each shows different strengths, suggesting substantial domain specialization and that reliable autonomous data engineering remains unresolved.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/30, 07:17 PMnot independentRepresentative
    DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness