Active Policy Optimization for Individualized Dosing via Gradient Variance Minimization
Abstract
In domains such as healthcare and marketing, learning optimal individualized dosing policies to maximize utility is crucial, yet high experimental costs impose strict budget constraints, necessitating efficient active policy learning. Existing active learning methods in causal inference primarily focus on binary treatments and effect estimation, leaving continuous dosing and policy optimization underexplored. To address this gap, we propose an active learning framework tailored for optimal policy learning. Exploiting the inherent structure of dose-response curves, we theoretically show that the policy optimization regret is bounded by the expected posterior gradient variance at the estimated optimal doses. Motivated by this result, we introduce Gradient Variance Active Learning for Individualized Dosing (GVALID), a batch acquisition strategy that greedily selects samples to minimize target gradient variance for efficient policy learning. Experiments demonstrate that GVALID achieves superior performance under strict budget constraints.
Lay Summary
Many real-world decision problems, from precision medicine to personalized services, require choosing the best action or dose for each individual, but testing many possible options can be expensive, slow, or impractical. This paper studies how to learn personalized dosing policies when only a limited number of experiments can be performed. Instead of trying to estimate the entire dose-response curve equally well everywhere, we focus on the information that matters most for making good decisions: whether the currently recommended dose is close to being optimal. We show that the error of the learned policy can be controlled by the uncertainty about how the response changes near the estimated best dose. Based on this idea, we propose GVALID, an active learning method that selects new experiments to reduce this key uncertainty as efficiently as possible. Experiments show that GVALID can learn better dosing policies than existing sampling strategies under tight budgets. This can help make experimental studies more sample-efficient and support better individualized decision-making when data collection is costly.