Robust Contextual Optimization with Missing Covariates
Abstract
Modern decision-making increasingly relies on contextual features (covariates) to improve optimization under uncertainty. In practice, however, such covariates are often only partially observed due to, e.g., data source heterogeneity or costly data collection. Nonetheless, most existing methods assume fully observed historical data and can become unreliable when this assumption is violated. We address this gap by proposing a distributionally robust optimization approach that exploits incomplete covariates to produce robust decisions without imputing a complete dataset. Our method builds ambiguity sets from the observed partial data and incorporates the general structure of the missingness mechanism, ensuring candidate distributions remain consistent with what is observed. Across settings with discrete or continuous covariates and outcomes, we derive tractable reformulations and establish finite-sample out-of-sample performance guarantees. Empirical results across a range of contextual decision-making tasks demonstrate that the proposed integrated approach consistently outperforms state-of-the-art baselines, including various impute-then-optimize pipelines, in both out-of-sample performance and reliability.
Lay Summary
How can we make reliable decisions when the features, or covariates, in a dataset are only partially observed? A common answer is to first preprocess the data by filling in the missing features, and then use the completed dataset for decision-making. However, we find that this widely used approach does not necessarily lead to good downstream performance or reliable decisions. The way features go missing can itself carry important information: If missing values are filled in before optimization, the downstream decision problem may no longer be able to use this structure. We propose a method that makes decisions directly from the partially observed data. Instead of completing the dataset first, our method keeps all plausible full-data distributions that are consistent with the observed records and the missingness process. It then chooses the decision that performs best against the most unfavorable distribution in this set. We show that this approach has finite-sample reliability guarantees and can be solved tractably in several settings. In experiments, it outperforms many preprocess-then-optimize methods, including pipelines that combine imputation with standard robust method, in both out-of-sample performance and reliability.