Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
Abstract
Foundation models are increasingly central to high-stakes AI systems, and governance frameworks now depend on evaluations to assess their risks and capabilities. Although general capability evaluations are widespread, social impact assessments covering bias, fairness, privacy, environmental costs, and labor remain uneven. To characterize this landscape, we conduct the first comprehensive analysis of social impact evaluation reporting, examining 186 first-party release reports and 248 third-party evaluation sources, supplemented by developer interviews. We find a stark division of labor: first-party reporting is sparse, often superficial, and declining in areas like environmental impact and bias, while third-party evaluators provide broader, more rigorous coverage of bias, harmful content, and performance disparities. However, only developers can authoritatively report on data provenance, content moderation labor, costs, and infrastructure, yet interviews reveal these disclosures are deprioritized unless tied to product adoption or compliance. Current practices leave major gaps in assessing societal impacts, underscoring the need for policies that mandate developer transparency, strengthen independent evaluation ecosystems, and create shared infrastructure for aggregating third-party evaluations.
Lay Summary
AI systems like ChatGPT now influence important decisions, and the public, regulators, and companies usually rely on the test results developers publish to judge whether a model is appropriate for a use case. But the social effects of these models, such as bias, privacy risks, environmental cost, and the working conditions of the people who label training data, get reported inconsistently or not at all. We did the first large-scale study of this, examining 186 reports developers released with their own models and 248 evaluations by outside researchers, plus interviews with people who build and test models. We found a clear split: developers' own reporting is thin, vague, and in some areas getting worse coverage over time, while independent researchers cover bias and harmful content more thoroughly. Yet some things only developers can report, like the labor behind content moderation, and companies tend to skip. These gaps leave real blind spots in how we understand AI's effects on society. We argue this calls for policies requiring more developer transparency, stronger support for independent evaluators, and shared tools for aggregating outside evaluations.