GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses
Abstract
Rather than replacing scientific judgment, we study how LLMs can support researchers through constructive feedback generation: producing targeted, actionable feedback that helps improve scientific papers. We operationalize feedback quality using two author-centric signals derived from author responses: validity, whether authors acknowledge feedback as correct, and actionability, whether they commit to follow-up actions. We introduce GoodPoint-ICLR, a dataset of 19K ICLR papers and review discussions annotated along these dimensions, and GoodPoint, a training recipe that fine-tunes on successful feedback and further aligns models using real and synthetic preference pairs. Across 1.2K ICLR papers, a GoodPoint-trained Qwen3-8B improves the predicted feedback success rate by 83.7% over its base model and achieves the strongest performance among size-comparable open models in matching human consensus feedback, surpassing Gemini-3-flash in precision. Human evaluation with paper authors further shows improvements in validity, actionability, specificity, and helpfulness.