LitReview Arena: Evaluating Literature Review Agents with Battle-style Peer Review Platform
Abstract
Lay Summary
Scientific literature reviews help researchers understand a field, compare existing work, and identify questions that are worth studying next. As AI systems become better at writing such reviews, we need to know whether their outputs are useful to researchers, rather than only checking whether they include the right keywords or citations. We built LitReview Arena, a platform where domain experts compare anonymized literature review drafts side by side and judge their coverage, evidence, organization, research suggestions, and overall usefulness. Based on these expert comparisons, we find that current AI systems still lag far behind human-written drafts, especially in parts of the review that require judgment and synthesis. We also find that automatic evaluators often disagree with expert researchers, so we develop LitJudge, an evaluator designed to better match expert preferences. Our work releases a public dataset and an evaluation process to help future AI systems produce scientific reviews that are more reliable and more useful to researchers.