Efficient Distributionally Robust Assortment Optimization in MNL Bandits
Abstract
Lay Summary
Online retailers and recommendation systems decide which products to show customers to maximize revenue. In practice, customer behavior often differs from what a model trained on historical data predicts: the underlying preferences may shift, or the model itself may be misspecified. We develop algorithms that choose product assortments that remain effective even when the customer's actual choice distribution lies within a plausible range around the learned model. We design efficient procedures under three standard ways of measuring distributional uncertainty, prove finite-sample guarantees on their performance, and validate them empirically. The result is a practical and theoretically grounded recipe for assortment decisions that hedge against realistic data drift.