Beyond Rational Illusion: Behaviorally Realistic Strategic Classification
Abstract
Strategic classification studies the interaction between decision models and agents who strategically manipulate their features for favorable outcomes. Existing SC frameworks typically rely on the idealized assumption that agents are strictly rational. However, evidence from behavioral economics and psychology consistently shows that real-world decision-making is often shaped by cognitive biases, deviating from pure rationality. To formalize this limitation, we identify and define a new problem setting, termed the behaviorally realistic strategic classification problem, where agents’ strategic manipulations deviate from full rationality due to psychological biases. Motivated by the identified limitation, we propose the Prospect-Guided Strategic Framework (Pro-SF) to address the problem, a principled framework grounded in prospect theory to model and learn under behaviorally realistic strategic responses. Specifically, to capture behaviorally realistic strategic manipulations, our framework reformulates the Stackelberg-style interaction between agents and the decision-maker by incorporating three key mechanisms inspired by prospect theory, including the asymmetry between benefits and costs, different subjective reference points, and non-rational probability distortion. Experiments on synthetic and real-world datasets establish Pro-SF as a behaviorally grounded approach to strategic classification, bridging machine learning and behavioral economics for more reliable deployment in the real world.
Lay Summary
When AI systems make decisions that affect people's lives---such as approving loans, screening job applicants, or assessing credit risk---people naturally try to present themselves in the best possible light. Existing AI safety research on this "gaming" behavior typically assumes that people act with perfect rationality: they calculate exactly what to change about themselves to get a favorable decision, weighing every cost and benefit precisely. But real people don't work that way. Someone applying for a loan might be so discouraged by how far they fall from the approval threshold that they don't even try to improve their profile. Another might overestimate their chances and make far more changes than necessary. These behaviors are driven by well-known psychological tendencies---like fearing losses more than valuing equivalent gains, anchoring judgments to personal reference points, and overestimating the likelihood of unlikely outcomes. We show that when AI classifiers are built assuming perfectly rational users, they systematically fail in practice: they either over-defend against gaming that never happens, or under-defend against gaming that goes further than expected. Both errors reduce accuracy for everyone. To fix this, we introduce Pro-SF, a framework that models how people actually behave when trying to game a system, drawing on prospect theory from behavioral economics. Pro-SF trains classifiers that anticipate psychologically realistic responses rather than idealized ones. Across multiple real-world datasets and human-subject experiments, Pro-SF substantially outperforms standard approaches, offering a more honest and reliable foundation for AI systems that interact with strategic human behavior.