Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Bowen, Ma, Haochen, Wang, Yuxin, Yang, Jie, Zheng, Yining, Chen, Xinchi, Huang, Xuanjing, Qiu, Xipeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911613793075200
author Li, Bowen
Ma, Haochen
Wang, Yuxin
Yang, Jie
Zheng, Yining
Chen, Xinchi
Huang, Xuanjing
Qiu, Xipeng
author_facet Li, Bowen
Ma, Haochen
Wang, Yuxin
Yang, Jie
Zheng, Yining
Chen, Xinchi
Huang, Xuanjing
Qiu, Xipeng
contents The rapid adoption of Large Language Models (LLMs) has spurred interest in automated peer review; however, progress is currently stifled by benchmarks that treat reviewing primarily as a rating prediction task. We argue that the utility of a review lies in its textual justification--its arguments, questions, and critique--rather than a scalar score. To address this, we introduce Beyond Rating, a holistic evaluation framework that assesses AI reviewers across five dimensions: Content Faithfulness, Argumentative Alignment, Focus Consistency, Question Constructiveness, and AI-Likelihood. Notably, we propose a Max-Recall strategy to accommodate valid expert disagreement and introduce a curated dataset of paper with high-confidence reviews, rigorously filtered to remove procedural noise. Extensive experiments demonstrate that while traditional n-gram metrics fail to reflect human preferences, our proposed text-centric metrics--particularly the recall of weakness arguments--correlate strongly with rating accuracy. These findings establish that aligning AI critique focus with human experts is a prerequisite for reliable automated scoring, offering a robust standard for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19502
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
Li, Bowen
Ma, Haochen
Wang, Yuxin
Yang, Jie
Zheng, Yining
Chen, Xinchi
Huang, Xuanjing
Qiu, Xipeng
Computation and Language
I.2.7; I.2.11
The rapid adoption of Large Language Models (LLMs) has spurred interest in automated peer review; however, progress is currently stifled by benchmarks that treat reviewing primarily as a rating prediction task. We argue that the utility of a review lies in its textual justification--its arguments, questions, and critique--rather than a scalar score. To address this, we introduce Beyond Rating, a holistic evaluation framework that assesses AI reviewers across five dimensions: Content Faithfulness, Argumentative Alignment, Focus Consistency, Question Constructiveness, and AI-Likelihood. Notably, we propose a Max-Recall strategy to accommodate valid expert disagreement and introduce a curated dataset of paper with high-confidence reviews, rigorously filtered to remove procedural noise. Extensive experiments demonstrate that while traditional n-gram metrics fail to reflect human preferences, our proposed text-centric metrics--particularly the recall of weakness arguments--correlate strongly with rating accuracy. These findings establish that aligning AI critique focus with human experts is a prerequisite for reliable automated scoring, offering a robust standard for future research.
title Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
topic Computation and Language
I.2.7; I.2.11
url https://arxiv.org/abs/2604.19502