Surrogate Metric Evaluation is a Causal Inference Problem
Abstract
Many AI systems are evaluated and trained using surrogate metrics that stand in for quantities of ultimate interest. Analyses of surrogate trustworthiness often focus on whether the surrogate and the target quantity are correlated, but this leaves open whether improvements in the surrogate will continue to improve the target under increasing optimization pressure. In my talk, I argue that this question can be conceptualized and addressed using dynamic causal models of feedback control systems. Using the classic example of the Watt governor, I illustrate how such systems exploit feedback loops to create higher-scale relationships by which one quantity adjusts to respond to another (e.g. the governor adjusts steam supply to match steam demand). Surrogate optimization similarly employs feedback in an attempt to make improvements in the surrogate correspond to improvements in the target, and optimization ceases to be useful once proposed improvements to the surrogate fail to match improvements in the target. It is this “matching” relationship, rather than the correlation between surrogate and target, that is relevant for analyzing optimization failures, reward hacking, and Goodhart-style effects.