DataClawEval introduces a benchmark for end-to-end data engineering agents using 100 production-grade tasks written by enterprise data engineers. The tasks span PySpark, MySQL, HiveSQL, PrestoSQL/Trino, and FlinkSQL, and are executed in isolated, case-specific sandboxes with deterministic rule-based grading rather than LLM judges. Across 16 frontier agents, the best overall score reaches only 74.9. No model dominates across all engines: each shows different strengths, suggesting substantial domain specialization and that reliable autonomous data engineering remains unresolved.
No heat snapshots are available in the last 24 hours.