Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Sihong, Jiang, Owen, Zhao, Yilun, Hu, Tiansheng, Ma, Yiling, Zhang, Kaiyan, Patwardhan, Manasi, Cohan, Arman |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation
by: Wu, Sihong, et al.
Published: (2026)
by: Wu, Sihong, et al.
Published: (2026)
SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature
by: Ding, Hang, et al.
Published: (2025)
by: Ding, Hang, et al.
Published: (2025)
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
by: Xu, Zhijian, et al.
Published: (2025)
by: Xu, Zhijian, et al.
Published: (2025)
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
by: Long, Yitao, et al.
Published: (2025)
by: Long, Yitao, et al.
Published: (2025)
Defend: Automated Rebuttals for Peer Review with Minimal Author Guidance
by: Khatri, Jyotsana, et al.
Published: (2026)
by: Khatri, Jyotsana, et al.
Published: (2026)
ResearchGym: Evaluating Language Model Agents on Real-World AI Research
by: Garikaparthi, Aniketh, et al.
Published: (2026)
by: Garikaparthi, Aniketh, et al.
Published: (2026)
SciMDR: Advancing Scientific Multimodal Document Reasoning
by: Chen, Ziyu, et al.
Published: (2026)
by: Chen, Ziyu, et al.
Published: (2026)
FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
by: Hu, Tiansheng, et al.
Published: (2025)
by: Hu, Tiansheng, et al.
Published: (2025)
IRIS: Interactive Research Ideation System for Accelerating Scientific Discovery
by: Garikaparthi, Aniketh, et al.
Published: (2025)
by: Garikaparthi, Aniketh, et al.
Published: (2025)
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research
by: Zhao, Yilun, et al.
Published: (2025)
by: Zhao, Yilun, et al.
Published: (2025)
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
by: Hu, Tiansheng, et al.
Published: (2026)
by: Hu, Tiansheng, et al.
Published: (2026)
PeerPrism: Peer Evaluation Expertise vs Review-writing AI
by: Sadeghian, Soroush, et al.
Published: (2026)
by: Sadeghian, Soroush, et al.
Published: (2026)
What Can Natural Language Processing Do for Peer Review?
by: Kuznetsov, Ilia, et al.
Published: (2024)
by: Kuznetsov, Ilia, et al.
Published: (2024)
MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search
by: Hu, Yunhai, et al.
Published: (2025)
by: Hu, Yunhai, et al.
Published: (2025)
MIR: Methodology Inspiration Retrieval for Scientific Research Problems
by: Garikaparthi, Aniketh, et al.
Published: (2025)
by: Garikaparthi, Aniketh, et al.
Published: (2025)
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
by: Zhao, Yilun, et al.
Published: (2025)
by: Zhao, Yilun, et al.
Published: (2025)
ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review
by: Sahu, Gaurav, et al.
Published: (2025)
by: Sahu, Gaurav, et al.
Published: (2025)
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
by: Yu, Zhaojian, et al.
Published: (2024)
by: Yu, Zhaojian, et al.
Published: (2024)
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
by: Wang, Chengye, et al.
Published: (2025)
by: Wang, Chengye, et al.
Published: (2025)
Identifying Aspects in Peer Reviews
by: Lu, Sheng, et al.
Published: (2025)
by: Lu, Sheng, et al.
Published: (2025)
The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
by: Sadallah, Abdelrahman, et al.
Published: (2025)
by: Sadallah, Abdelrahman, et al.
Published: (2025)
SUCEA: Reasoning-Intensive Retrieval for Adversarial Fact-checking through Claim Decomposition and Editing
by: Liu, Hongjun, et al.
Published: (2025)
by: Liu, Hongjun, et al.
Published: (2025)
Table-R1: Inference-Time Scaling for Table Reasoning
by: Yang, Zheyuan, et al.
Published: (2025)
by: Yang, Zheyuan, et al.
Published: (2025)
PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing
by: Żurawicki, Krzysztof, et al.
Published: (2026)
by: Żurawicki, Krzysztof, et al.
Published: (2026)
Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?
by: Tang, Xiangru, et al.
Published: (2023)
by: Tang, Xiangru, et al.
Published: (2023)
AgentReview: Exploring Peer Review Dynamics with LLM Agents
by: Jin, Yiqiao, et al.
Published: (2024)
by: Jin, Yiqiao, et al.
Published: (2024)
Is Peer-Reviewing Worth the Effort?
by: Church, Kenneth, et al.
Published: (2024)
by: Church, Kenneth, et al.
Published: (2024)
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
by: Zhao, Yilun, et al.
Published: (2025)
by: Zhao, Yilun, et al.
Published: (2025)
Disparities in Peer Review Tone and the Role of Reviewer Anonymity
by: Sahakyan, Maria, et al.
Published: (2025)
by: Sahakyan, Maria, et al.
Published: (2025)
REVERE: Reflective Evolving Research Engineer for Scientific Workflows
by: Gangireddi, Balaji Dinesh, et al.
Published: (2026)
by: Gangireddi, Balaji Dinesh, et al.
Published: (2026)
Automated Peer Reviewing in Paper SEA: Standardization, Evaluation, and Analysis
by: Yu, Jianxiang, et al.
Published: (2024)
by: Yu, Jianxiang, et al.
Published: (2024)
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
by: Zhao, Yilun, et al.
Published: (2026)
by: Zhao, Yilun, et al.
Published: (2026)
Z1: Efficient Test-time Scaling with Code
by: Yu, Zhaojian, et al.
Published: (2025)
by: Yu, Zhaojian, et al.
Published: (2025)
Detecting AI-Generated Content in Academic Peer Reviews
by: Shen, Siyuan, et al.
Published: (2026)
by: Shen, Siyuan, et al.
Published: (2026)
PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers
by: Loc, Ngoc Phan Phuoc, et al.
Published: (2026)
by: Loc, Ngoc Phan Phuoc, et al.
Published: (2026)
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
by: Yu, Sungduk, et al.
Published: (2025)
by: Yu, Sungduk, et al.
Published: (2025)
Is Your Paper Being Reviewed by an LLM? Investigating AI Text Detectability in Peer Review
by: Yu, Sungduk, et al.
Published: (2024)
by: Yu, Sungduk, et al.
Published: (2024)
RATE: Reviewer Profiling and Annotation-free Training for Expertise Ranking in Peer Review Systems
by: Liu, Weicong, et al.
Published: (2026)
by: Liu, Weicong, et al.
Published: (2026)
PeerQA: A Scientific Question Answering Dataset from Peer Reviews
by: Baumgärtner, Tim, et al.
Published: (2025)
by: Baumgärtner, Tim, et al.
Published: (2025)
Survey on Evaluation of LLM-based Agents
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
Similar Items
-
RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation
by: Wu, Sihong, et al.
Published: (2026) -
SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature
by: Ding, Hang, et al.
Published: (2025) -
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
by: Xu, Zhijian, et al.
Published: (2025) -
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
by: Long, Yitao, et al.
Published: (2025) -
Defend: Automated Rebuttals for Peer Review with Minimal Author Guidance
by: Khatri, Jyotsana, et al.
Published: (2026)