Personalized Policy Learning through Discrete Experimentation
Zhiqi Zhang ⋅ Zhiyu Zeng ⋅ Ruohan Zhan ⋅ Dennis Zhang
Abstract
While Randomized controlled trials (RCTs), or A/B tests, are the gold standard for optimizing online-platform policies, they are limited by discrete testing levels. This approach is suboptimal for continuous variables (e.g., prices and incentives), as it fails to extrapolate to untested values or account for user heterogeneity. We address this by developing Deep Learning for Policy Targeting (\textsf{DLPT}) to learn personalized continuous policies from discrete RCTs using high-dimensional features. We prove our estimators are asymptotically unbiased and consistent, achieving a $\sqrt{n}$-regret bound. In a collaboration with a leading social media platform to optimize creator incentives, we show that \textsf{DLPT} substantially outperforms existing benchmarks.
Lay Summary
Tech companies constantly use A/B testing to figure out what works best, such as offering users a $5 versus a $10 reward. However, these tests are fundamentally limited because they only check a few fixed options. They cannot easily predict what would happen at untested amounts, like $7.50, nor can they intuitively personalize these amounts for users with different habits. To solve this, we developed a new AI method called Deep Learning for Policy Targeting (DLPT). Our approach takes the restricted data from standard A/B tests and uses machine learning to smoothly "fill in the blanks" for all untested values. More importantly, it analyzes individual user characteristics to pinpoint the exact, personalized reward that works best for each person. We deployed DLPT on a major social media platform to optimize the financial incentives given to content creators, where it drastically outperformed existing methods. Ultimately, our research empowers organizations to make smarter, highly personalized decisions from their existing data, without needing to run complex or costly new experiments.
Successful Page Load