Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

CausalDS: Benchmarking Causal Reasoning in Data-Science Agents

First seen · 7/10/2026, 12:00 PMLatest activity · 7/10/2026, 12:00 PM

CausalDS is a benchmark for evaluating causal reasoning in agentic data-science workflows. Each instance combines a sampled structural causal model, generated observational data, and a synthetic natural-language story grounded in a realistic domain. Tasks span all three of Pearl’s causal rungs, from prediction to intervention and counterfactual reasoning. The benchmark also evaluates coding and tool use, imperfect observations, uncertainty quantification, and abstention when the evidence does not warrant an answer. Its synthetic generation aims to preserve empirical structure while reducing memorization of recurring causal examples.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/10, 12:00 PMnot independentRepresentative
    CausalDS: Benchmarking Causal Reasoning in Data-Science Agents