arXivZi-Rong Li
CausalArena: Benchmarking Causal Discovery in the Foundation Model Era
Papers75
As causal discovery foundation models emerge, static benchmarks struggle to separate genuine structural reasoning from pretraining overlap. CausalArena introduces an evolving evaluation suite spanning synthetic graphs, semantic operational environments, formula-grounded scientific mechanisms, and real-world datasets. Experiments across classical, neural, and pretrained methods show substantial ranking shifts between regimes, highlighting the fragility of performance claims and the necessity of multi-environment testing.
Why it's worth reading
As foundation models enter causal inference, this benchmark exposes how severe ranking instability and pretraining overlap challenge prevailing evaluation practices.
Tags
因果推断因果发现基准评测Foundation ModelsCausalArena机器学习