Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
Abstract
A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polishing, paraphrasing, and grammar correction of otherwise human-written reviews. But, are these policies enforceable? To answer this question, we assemble a dataset of peer reviews simulating multiple levels of human-AI collaboration, and evaluate five state-of-the-art detectors, including two commercial systems. Our analysis shows that all detectors misclassify a non-trivial fraction of LLM-polished reviews as AI-generated, thereby risking false accusations of academic misconduct. We further investigate whether peer-review-specific signals, including access to the paper manuscript and the constrained domain of scientific writing, can be leveraged to improve detection. While incorporating such signals yields measurable gains in some settings, we identify limitations in each approach and find that none meets the accuracy standards required for identifying AI use in peer reviews. Importantly, our results suggest that recent public estimates of AI use in peer reviews through the use of current AI-text detectors should be interpreted with caution, as they misclassify mixed reviews (collaborative human-AI outputs) as fully AI generated, potentially overstating the extent of policy violations.
Lay Summary
Scientific research undergoes rigorous evaluation by expert human reviewers before they are published in conference and journal proceedings. This process, called peer review, is essential in ensuring the quality and integrity of published research. Lately, reviewers are turning to AI to assist them in writing these reviews. But, conference organizers want human judgement, not AI judgement, in reviews. This has led them to enact policies that allow reviewers to use AI to polish the prose (fix grammar, fluency etc), but not write the substance. AI detectors seem to be an obvious tool to enforce these rules. We find that today's AI detectors are not good enough to reliably tell AI-polished text from fully AI-generated text. In large scale conferences like ICML and NeurIPS, the best AI detectors can misclassify thousands of policy-abiding reviews, putting their authors at risk of reputational damage and bans from the very conferences they serve. This means, if publishing venues are looking to enforce polishing-only policies, AI detectors are not the answer. It also means public estimates about how many reviews were AI-written, including figures that got significant media attention, are likely inflated. These findings likely extend beyond peer review, anywhere AI detection is used to enforce policy, whether in education, journalism, or social media moderation.