When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhu, Changjia, Xiong, Junjie, Ma, Renkai, Lu, Zhicong, Liu, Yao, Li, Lingyao
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916947288915968
author Zhu, Changjia
Xiong, Junjie
Ma, Renkai
Lu, Zhicong
Liu, Yao
Li, Lingyao
author_facet Zhu, Changjia
Xiong, Junjie
Ma, Renkai
Lu, Zhicong
Liu, Yao
Li, Lingyao
contents Peer review is the cornerstone of academic publishing, yet the process is increasingly strained by rising submission volumes, reviewer overload, and expertise mismatches. Large language models (LLMs) are now being used as "reviewer aids," raising concerns about their fairness, consistency, and robustness against indirect prompt injection attacks. This paper presents a systematic evaluation of LLMs as academic reviewers. Using a curated dataset of 1,441 papers from ICLR 2023 and NeurIPS 2022, we evaluate GPT-5-mini against human reviewers across ratings, strengths, and weaknesses. The evaluation employs structured prompting with reference paper calibration, topic modeling, and similarity analysis to compare review content. We further embed covert instructions into PDF submissions to assess LLMs' susceptibility to prompt injection. Our findings show that LLMs consistently inflate ratings for weaker papers while aligning more closely with human judgments on stronger contributions. Moreover, while overarching malicious prompts induce only minor shifts in topical focus, explicitly field-specific instructions successfully manipulate specific aspects of LLM-generated reviews. This study underscores both the promises and perils of integrating LLMs into peer review and points to the importance of designing safeguards that ensure integrity and trust in future review processes.
format Preprint
id arxiv_https___arxiv_org_abs_2509_09912
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
Zhu, Changjia
Xiong, Junjie
Ma, Renkai
Lu, Zhicong
Liu, Yao
Li, Lingyao
Computers and Society
Cryptography and Security
Peer review is the cornerstone of academic publishing, yet the process is increasingly strained by rising submission volumes, reviewer overload, and expertise mismatches. Large language models (LLMs) are now being used as "reviewer aids," raising concerns about their fairness, consistency, and robustness against indirect prompt injection attacks. This paper presents a systematic evaluation of LLMs as academic reviewers. Using a curated dataset of 1,441 papers from ICLR 2023 and NeurIPS 2022, we evaluate GPT-5-mini against human reviewers across ratings, strengths, and weaknesses. The evaluation employs structured prompting with reference paper calibration, topic modeling, and similarity analysis to compare review content. We further embed covert instructions into PDF submissions to assess LLMs' susceptibility to prompt injection. Our findings show that LLMs consistently inflate ratings for weaker papers while aligning more closely with human judgments on stronger contributions. Moreover, while overarching malicious prompts induce only minor shifts in topical focus, explicitly field-specific instructions successfully manipulate specific aspects of LLM-generated reviews. This study underscores both the promises and perils of integrating LLMs into peer review and points to the importance of designing safeguards that ensure integrity and trust in future review processes.
title When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
topic Computers and Society
Cryptography and Security
url https://arxiv.org/abs/2509.09912