The paper introduces CALVER, a training-free symbolic verifier for selecting among sampled LLM causal-reasoning traces. Instead of voting on final answers, it checks structured candidates against Pearl-style causal criteria such as d-separation, backdoor adjustment, and intervention. On frozen candidate pools for CLEAR queries with multiple graph-valid answers, CALVER achieves 42.1%, while plurality voting, a reward model, an LLM judge, and model confidence remain near 30%. The reported gains extend across ten published Bayesian networks, another model family, text-derived graphs, treatment-effect decisions, and a logic task checked with truth tables.
No heat snapshots are available in the last 24 hours.