Desirable Effort Fairness and Optimality Trade-offs in Strategic Learning
Abstract
Strategic classification examines how decision rules interact with agents who strategically adapt their features. Most existing models focus on maximizing predictive performance, assuming agents best respond to the learned classifier. However, real decision-making systems are rarely optimized solely for accuracy: ethical, economic, and institutional considerations often make some feature changes more desirable than others. At the same time, principals may wish to incentivize these changes fairly across heterogeneous agents. While prior work has studied causal structure between features, notions of desirability, and information disparities in isolation, this work initiates a unified treatment of these components within a single framework. We frame the problem as a constrained optimization problem that captures the trade-offs between optimality, desirability, and fairness. We provide theoretical guarantees on the principal's optimality loss constrained to a particular desirability fairness tolerance for multiple broad classes of fairness measures. Finally, through experiments on real datasets, we show the explicit tradeoff between maximizing accuracy and fairness in desirability effort.
Lay Summary
Many decisions about people, such as loan approvals, recommendations, or admissions, are increasingly influenced by automated systems. People who are affected by these systems may change their behavior in response: for example, they may try to improve the features that the system rewards. However, not all changes are equally valuable. Some changes may reflect genuine improvement, while others may simply be attempts to game the system. Moreover, different groups of people may have different information, costs, or opportunities, so the same system may create unequal incentives across groups. This paper studies how a decision-maker can encourage useful changes while treating different groups fairly. We build a model in which people respond to a decision rule, but may only have partial information about it, and where changing one feature can also affect other features. We then ask how much performance the decision-maker must give up in order to limit unfair differences in the incentives given to different groups. Our results give mathematical guarantees on this trade-off for broad classes of fairness requirements. These guarantees can help a decision-maker understand, before deploying a system, how strict fairness requirements may affect goals such as accuracy or overall benefit. We also run experiments on real datasets showing how the fairness-performance trade-off depends on the way group differences align with the features considered desirable to improve.