A Theory of Data Acquisition and Pricing at Scale
Abstract
Data plays an invaluable role in large-scale ML training pipelines. Multiple factors, including the need to incentivize the creation of high-quality data and efforts to compensate creative data work, have led to increased interest in data pricing. Data pricing mechanisms seek to establish a market where data providers are compensated based (in part) on the value of their data to the data buyer, e.g., frontier AI labs. However, assessing the exact value that each provider's data adds to the data buyer's objective requires repeated re-training, which is infeasible in practice. Our work studies data pricing under compute constraints. In our setting, data buyers cannot make data acquisition decisions optimally due to limited compute. Inspired by existing practice in the field of data selection, we propose a model for this problem called ``pricing with an attribution oracle,'' and provide a theoretical analysis of compute-efficient acquisition and pricing.
Lay Summary
High-quality data is valuable for training machine learning models, but deciding how much each data provider should be paid is difficult: the most direct way to measure a dataset's value would be to repeatedly retrain the model, which is far too expensive at modern scale. This paper studies how a data buyer can make pricing and purchasing decisions when computation is limited. We show that limited compute changes both which data should be bought and whether sellers have incentives to report their true costs. Our results give a theoretical framework for designing data-buying rules that account for computation while still supporting truthful payments.