CausalDS is a benchmark for evaluating causal reasoning in agentic data-science workflows. Each instance combines a sampled structural causal model, generated observational data, and a synthetic natural-language story grounded in a realistic domain. Tasks span all three of Pearl’s causal rungs, from prediction to intervention and counterfactual reasoning. The benchmark also evaluates coding and tool use, imperfect observations, uncertainty quantification, and abstention when the evidence does not warrant an answer. Its synthetic generation aims to preserve empirical structure while reducing memorization of recurring causal examples.
No heat snapshots are available in the last 24 hours.