Partial Identification of Policy Values under Network Interference
Abstract
Offline Policy Evaluation (OPE) aims to estimate the value of a target policy from historical logged data without interating with the environment, thereby assessing policy performance. In settings with network interference, individuals no longer satisfy the SUTVA assumption: an individual’s outcome is influenced not only by their own treatment but also by the treatments of their neighbors, which makes the definition and estimation of policy value more complex. To capture this interference mechanism, we allow all neighbors to affect individual outcomes through a unified exposure mapping, and we use a decaying higher-order neighborhood aggregation to characterize the influence of more distant neighbors. Moreover, in real-world applications, the target policy and the logging policy often do not fully overlap (non-overlap), so the policy value in non-overlap regions cannot be point-identified. To address this issue, we partially identify the policy value over non-overlap regions and, under a smoothness assumption, formulate the estimation of the lower and upper bounds as a linear program, yielding valid bounds on the offline policy value. Finally, we conduct systematic experiments on semi-synthetic network data to validate the effectiveness and robustness of the proposed method under network interference and limited overlap.
Lay Summary
We study how to evaluate new decision-making rules (policies) using only historical data, without actually testing them in the real world. In complex networks—like social networks or transportation systems—one person’s outcome can be affected not just by their own treatment but also by what happens to their neighbors. This makes it tricky to measure how well a policy would work. To tackle this, we consider all neighbors’ influence and use a method that gradually reduces the effect of more distant neighbors. Sometimes, the new policy targets situations that the historical data didn’t cover, so we can’t determine exact results for those cases. Instead, we estimate safe upper and lower limits for the policy’s impact using a mathematical approach called a linear program, assuming outcomes change smoothly across similar situations. We tested our approach on simulated network data and found that it reliably measures policy performance even under network interference and limited data coverage. This method helps researchers confidently assess policies in realistic networked settings without running risky or costly real-world experiments.