AgenticDataBench is proposed as a broad benchmark for evaluating LLM-based data agents across realistic data-science workflows. It covers 15 vertical domains, including five real-world B2B use cases from a fintech company. The benchmark organizes recurring data-centric operations as reusable skills, extracts representative skills through skill-aligned hierarchical clustering of Stack Overflow solutions, and selects or generates tasks to improve coverage. The authors also provide fine-grained ground-truth annotations, skill-level evaluation insights for state-of-the-art data agents, and an open-source testbed.
No heat snapshots are available in the last 24 hours.