The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sadallah, Abdelrahman, Baumgärtner, Tim, Gurevych, Iryna, Briscoe, Ted
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914050243297280
author Sadallah, Abdelrahman
Baumgärtner, Tim
Gurevych, Iryna
Briscoe, Ted
author_facet Sadallah, Abdelrahman
Baumgärtner, Tim
Gurevych, Iryna
Briscoe, Ted
contents Providing constructive feedback to paper authors is a core component of peer review. With reviewers increasingly having less time to perform reviews, automated support systems are required to ensure high reviewing quality, thus making the feedback in reviews useful for authors. To this end, we identify four key aspects of review comments (individual points in weakness sections of reviews) that drive the utility for authors: Actionability, Grounding & Specificity, Verifiability, and Helpfulness. To enable evaluation and development of models assessing review comments, we introduce the RevUtil dataset. We collect 1,430 human-labeled review comments and scale our data with 10k synthetically labeled comments for training purposes. The synthetic data additionally contains rationales, i.e., explanations for the aspect score of a review comment. Employing the RevUtil dataset, we benchmark fine-tuned models for assessing review comments on these aspects and generating rationales. Our experiments demonstrate that these fine-tuned models achieve agreement levels with humans comparable to, and in some cases exceeding, those of powerful closed models like GPT-4o. Our analysis further reveals that machine-generated reviews generally underperform human reviews on our four aspects.
format Preprint
id arxiv_https___arxiv_org_abs_2509_04484
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
Sadallah, Abdelrahman
Baumgärtner, Tim
Gurevych, Iryna
Briscoe, Ted
Computation and Language
Artificial Intelligence
Computers and Society
Providing constructive feedback to paper authors is a core component of peer review. With reviewers increasingly having less time to perform reviews, automated support systems are required to ensure high reviewing quality, thus making the feedback in reviews useful for authors. To this end, we identify four key aspects of review comments (individual points in weakness sections of reviews) that drive the utility for authors: Actionability, Grounding & Specificity, Verifiability, and Helpfulness. To enable evaluation and development of models assessing review comments, we introduce the RevUtil dataset. We collect 1,430 human-labeled review comments and scale our data with 10k synthetically labeled comments for training purposes. The synthetic data additionally contains rationales, i.e., explanations for the aspect score of a review comment. Employing the RevUtil dataset, we benchmark fine-tuned models for assessing review comments on these aspects and generating rationales. Our experiments demonstrate that these fine-tuned models achieve agreement levels with humans comparable to, and in some cases exceeding, those of powerful closed models like GPT-4o. Our analysis further reveals that machine-generated reviews generally underperform human reviews on our four aspects.
title The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2509.04484