Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911613793075200 |
|---|---|
| author | Li, Bowen Ma, Haochen Wang, Yuxin Yang, Jie Zheng, Yining Chen, Xinchi Huang, Xuanjing Qiu, Xipeng |
| author_facet | Li, Bowen Ma, Haochen Wang, Yuxin Yang, Jie Zheng, Yining Chen, Xinchi Huang, Xuanjing Qiu, Xipeng |
| contents | The rapid adoption of Large Language Models (LLMs) has spurred interest in automated peer review; however, progress is currently stifled by benchmarks that treat reviewing primarily as a rating prediction task. We argue that the utility of a review lies in its textual justification--its arguments, questions, and critique--rather than a scalar score. To address this, we introduce Beyond Rating, a holistic evaluation framework that assesses AI reviewers across five dimensions: Content Faithfulness, Argumentative Alignment, Focus Consistency, Question Constructiveness, and AI-Likelihood. Notably, we propose a Max-Recall strategy to accommodate valid expert disagreement and introduce a curated dataset of paper with high-confidence reviews, rigorously filtered to remove procedural noise. Extensive experiments demonstrate that while traditional n-gram metrics fail to reflect human preferences, our proposed text-centric metrics--particularly the recall of weakness arguments--correlate strongly with rating accuracy. These findings establish that aligning AI critique focus with human experts is a prerequisite for reliable automated scoring, offering a robust standard for future research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_19502 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews Li, Bowen Ma, Haochen Wang, Yuxin Yang, Jie Zheng, Yining Chen, Xinchi Huang, Xuanjing Qiu, Xipeng Computation and Language I.2.7; I.2.11 The rapid adoption of Large Language Models (LLMs) has spurred interest in automated peer review; however, progress is currently stifled by benchmarks that treat reviewing primarily as a rating prediction task. We argue that the utility of a review lies in its textual justification--its arguments, questions, and critique--rather than a scalar score. To address this, we introduce Beyond Rating, a holistic evaluation framework that assesses AI reviewers across five dimensions: Content Faithfulness, Argumentative Alignment, Focus Consistency, Question Constructiveness, and AI-Likelihood. Notably, we propose a Max-Recall strategy to accommodate valid expert disagreement and introduce a curated dataset of paper with high-confidence reviews, rigorously filtered to remove procedural noise. Extensive experiments demonstrate that while traditional n-gram metrics fail to reflect human preferences, our proposed text-centric metrics--particularly the recall of weakness arguments--correlate strongly with rating accuracy. These findings establish that aligning AI critique focus with human experts is a prerequisite for reliable automated scoring, offering a robust standard for future research. |
| title | Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews |
| topic | Computation and Language I.2.7; I.2.11 |
| url | https://arxiv.org/abs/2604.19502 |