Sequential Conditional Independence Testing with Machine Learning Models
Abstract
Conditional Independence Testing is a ubiquitous problem in scientific discovery. The widely used Model-X assumption consists in shifting the modelling burden from the dependency of the output given the inputs, to the dependencies within inputs. Log-optimal e-variables have been studied in this setting, but it remains unclear how to include machine learning models in these tests. In this work, we exploit the performance drop of a model when a given feature is removed to construct a coin-betting e-value for bounded losses and an exponential e-value that mimics density-based approaches. Surprisingly, in a misspecified setting, we show theoretically and experimentally that neither e-value uniformly outperforms the other. Finally, we provide actionable algorithms for their implementation.