Incentivized Exploration with Stochastic Covariates: A Two-Stage Mechanism Design for Recommender System
Abstract
Lay Summary
(Problem) Recommendation systems must occasionally try unfamiliar options to discover what works best for each user, but users act in their own interest—they will ignore a suggestion that does not appear beneficial to them, making experimentation ineffective. This tension between learning and user trust is especially hard when every user is different and arrives with unique characteristics the system must handle on the fly. (Solution) We develop a two-stage algorithm that first explores new options while guaranteeing every recommendation remains credibly in the user's best interest, and then efficiently focuses on the most uncertain options to rapidly improve recommendation quality. Our mathematical analysis shows the algorithm's mistakes diminish over time and reveals a precise tradeoff: a larger persuasion budget accelerates learning but with diminishing returns. (Impact) This framework enables recommendation systems—including safety-critical applications like personalized medicine—to learn effectively without misleading users, building both trust and performance over time.