ReviewAgents: Bridging the Gap Between Human and AI-Generated Paper Reviews

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Xian, Ruan, Jiacheng, Zhang, Zongyun, Gao, Jingsheng, Liu, Ting, Fu, Yuzhuo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909690981515264
author Gao, Xian
Ruan, Jiacheng
Zhang, Zongyun
Gao, Jingsheng
Liu, Ting
Fu, Yuzhuo
author_facet Gao, Xian
Ruan, Jiacheng
Zhang, Zongyun
Gao, Jingsheng
Liu, Ting
Fu, Yuzhuo
contents Academic paper review is a critical yet time-consuming task within the research community. With the increasing volume of academic publications, automating the review process has become a significant challenge. The primary issue lies in generating comprehensive, accurate, and reasoning-consistent review comments that align with human reviewers' judgments. In this paper, we address this challenge by proposing ReviewAgents, a framework that leverages large language models (LLMs) to generate academic paper reviews. We first introduce a novel dataset, Review-CoT, consisting of 142k review comments, designed for training LLM agents. This dataset emulates the structured reasoning process of human reviewers-summarizing the paper, referencing relevant works, identifying strengths and weaknesses, and generating a review conclusion. Building upon this, we train LLM reviewer agents capable of structured reasoning using a relevant-paper-aware training method. Furthermore, we construct ReviewAgents, a multi-role, multi-LLM agent review framework, to enhance the review comment generation process. Additionally, we propose ReviewBench, a benchmark for evaluating the review comments generated by LLMs. Our experimental results on ReviewBench demonstrate that while existing LLMs exhibit a certain degree of potential for automating the review process, there remains a gap when compared to human-generated reviews. Moreover, our ReviewAgents framework further narrows this gap, outperforming advanced LLMs in generating review comments.
format Preprint
id arxiv_https___arxiv_org_abs_2503_08506
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReviewAgents: Bridging the Gap Between Human and AI-Generated Paper Reviews
Gao, Xian
Ruan, Jiacheng
Zhang, Zongyun
Gao, Jingsheng
Liu, Ting
Fu, Yuzhuo
Computation and Language
Academic paper review is a critical yet time-consuming task within the research community. With the increasing volume of academic publications, automating the review process has become a significant challenge. The primary issue lies in generating comprehensive, accurate, and reasoning-consistent review comments that align with human reviewers' judgments. In this paper, we address this challenge by proposing ReviewAgents, a framework that leverages large language models (LLMs) to generate academic paper reviews. We first introduce a novel dataset, Review-CoT, consisting of 142k review comments, designed for training LLM agents. This dataset emulates the structured reasoning process of human reviewers-summarizing the paper, referencing relevant works, identifying strengths and weaknesses, and generating a review conclusion. Building upon this, we train LLM reviewer agents capable of structured reasoning using a relevant-paper-aware training method. Furthermore, we construct ReviewAgents, a multi-role, multi-LLM agent review framework, to enhance the review comment generation process. Additionally, we propose ReviewBench, a benchmark for evaluating the review comments generated by LLMs. Our experimental results on ReviewBench demonstrate that while existing LLMs exhibit a certain degree of potential for automating the review process, there remains a gap when compared to human-generated reviews. Moreover, our ReviewAgents framework further narrows this gap, outperforming advanced LLMs in generating review comments.
title ReviewAgents: Bridging the Gap Between Human and AI-Generated Paper Reviews
topic Computation and Language
url https://arxiv.org/abs/2503.08506