Evaluating Bivariate Causal Statements Based on Mutual Compatibility
Abstract
For many real-world systems, causal ground truth is difficult to obtain, making claims about causal effects hard to assess. We develop methods for evaluating collections of bivariate causal statements, one for each pair of variables in a fixed system. In the setting of acyclic linear statements, any such collection can be extended to a unique multivariate causal model, but we argue that this induced model is implausible if it imposes substantial additional confounding to explain observed correlations. We introduce a compatibility score that quantifies this notion of plausibility, notably without relying on the faithfulness assumption. Additionally, we define an incompatibility score for purely graphical bivariate causal statements, based on global consistency constraints that are derived from acyclicity and faithfulness assumptions. We give theoretical and empirical evidence that both scores can successfully distinguish correct from incorrect causal statements in generic settings. Moreover, we demonstrate the practical applicability of our methods by analyzing causal claims made by large language models. Our work aims to provide a foundation for assessing the reliability of causal information derived from human experts or artificial intelligence in settings where alternative forms of validation are unavailable.
Lay Summary
Every day, we encounter causal claims that are hard to fact-check, especially if they have not been validated by scientific experiments. While it is difficult to assess a causal claim by itself, a collection of causal claims about the same system may reveal some internal inconsistencies. We developed an approach to check whether many pairwise causal claims can fit together as part of an overall causal model that satisfies a certain plausibility criterion. On this basis, we define compatibility scores that can be computed for causal claims in specific quantitative or qualitative formats, given that there are sufficiently many. These scores serve as basic sanity checks: a good score does not prove that the claims are correct, but a bad score indicates incompatibility and therefore provides evidence against them. We hope that our scores can help flag incorrect causal claims and improve the reliability of causal information in settings where other forms of validation are infeasible.