Articulate Intuition or Genuine Analysis? Benchmarking Epistemic Reliability in LLM-as-a-Judge Peer Review
Abstract
When an LLM judge calls a peer review analytical and a human committee calls another review high quality, are they tracking the same thing? We argue they are not, and that the difference matters philosophically. We operationalise Kahneman's dual-process theory into a structured rubric for peer review and release Kahneman4Review, a benchmark of 3,563 rated reviews scored along nine epistemic dimensions, eight bias diagnostics, and a continuous reasoning-quality score. Three findings bear on trustworthiness: decision tier is not detectably aligned with rubric-level epistemic quality; an agentic AI reviewer scores higher than humans, but most of the gap is explained after controlling for review length and venue; and ICLR review reasoning shows a sharp temporal discontinuity around the ChatGPT-deployment boundary that survives length-adjustment, while 2021 and 2022 baselines remain statistically indistinguishable. A matched function-probe pilot further shows the rubric can discriminate genuine fault-finding from surface fluency. We argue that a trustworthy reliability benchmark for LLM judges must separate analytical form from epistemic function, and propose concrete design choices toward that goal. An interactive demo is available at https://huggingface.co/spaces/anonymous-D1C4/space-D1C4.