ScrapeBench: Evaluating Legal Compliance of AI Agents in Website Scraping
Joseph Marvin Imperial ⋅ Daniel Slate ⋅ Noam Kolt
Abstract
AI agents deployed on the open web are becoming more capable of performing actions that may breach contracts and violate computer security laws. In this work, we introduce ScrapeBench, the first agentic benchmark for evaluating the legal compliance of frontier AI agents with website scraping policies, including terms of service (ToS) and robots.txt directives. Across widely used closed-weight and open-weight AI agents, we find that most agents exhibit low compliance rates (< 5\%) when instructed to scrape websites in our sample of 1,626 popular websites that explicitly prohibit scraping. In addition, we observe that even where AI agents actively check the applicable ToS and robots.txt, they continue to exhibit low compliance rates. Conversely, in our sample of 269 websites that do not prohibit scraping, we observe that agents regularly refuse to scrape despite it being lawful, such as Claude Opus 4.6 exhibiting an over-refusal rate of 98.8\%. Lastly, our human baseline experiments ($n$ = 180 participants) suggest that humans exhibit compliance rates comparable to AI agents but over-refuse at much higher rates. We release our anonymized code and data in this repository: https://anonymous.4open.science/r/scrapebench-75D6
Chat is not available.
Successful Page Load