A Minimax Approach for Optimal Intervention Policy Learning with Two-Stage Outcomes
Abstract
When designing interventions to promote desired actions, two-stage agent heterogeneity -- encompassing both engagement with the intervention and completion of the desired action -- creates significant challenges in identifying optimal intervention policies. While this two-dimensional heterogeneity creates distinct agent response types with varying marginal policy returns, existing literature typically falls short in full identification of all agent types, leading to inefficient intervention allocations. To address the challenge of learning optimal policies that account for two-stage outcomes, we propose a minimax approach within a counterfactual principal strata framework. A value function, accommodating varying policy returns across six potentially non-identifiable principal strata, is designed and partially identified to minimize the worst-case value loss relative to three benchmark policies: never-treat, always-treat, and oracle. We introduce three estimators for optimal policy learning: Principal Outcome Regression (P-OR), Principal Inverse Propensity Scoring (P-IPS), and Principal Doubly Robust (P-DR), providing theoretical guarantees for their unbiasedness, robustness, and regret upper bounds. Extensive numerical experiments demonstrate the effectiveness and superiority of the proposed approach.
Lay Summary
When designing interventions to promote desired actions, such as promotional incentives in e-commerce to increase purchases, educational programs to improve student outcomes, or healthcare reminders to boost treatment adherence, policymakers must account for agent-level heterogeneity in both engagement with the intervention itself and completion of the target action. This two-dimensional heterogeneity creates distinct agent types, or principal strata, yielding different returns for the policy. The fundamental challenge in learning an optimal policy here is non-identifiability of different response types. Because it is impossible to simultaneously observe how the same individual would have responded under both with and without intervention conditions, observational data alone cannot fully distinguish between different strata. To address this challenge, existing policy learning methods impose exclusion assumptions that posit the absence of certain principal strata. In this article, we propose a minimax approach within a counterfactual principal strata framework. Since the exact policy return (value function) is non-identifiable, our method utilizes partial identification to establish worst-case bounds. It derives the optimal policy by minimizing the worst-case value loss relative to three alternative policies (never treat, always treat, and oracle). Furthermore, we introduce three robust estimators, Principal Outcome Regression (P-OR), Principal Inverse Propensity Scoring (P-IPS), and Principal Doubly Robust (P-DR), providing solid theoretical guarantees for regret upper bounds. Extensive numerical experiments demonstrate the effectiveness and superiority of the proposed approach. This framework ensures reliable and efficient intervention targeting under data uncertainty, with broad applicability across digital platforms, education, and public policy.