LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Lingyao, Xiong, Junjie, Zhu, Changjia, Yu, Runlong, Chen, Chen, Wang, Junyu, Ma, Renkai, Lu, Zhicong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914599120404480
author Li, Lingyao
Xiong, Junjie
Zhu, Changjia
Yu, Runlong
Chen, Chen
Wang, Junyu
Ma, Renkai
Lu, Zhicong
author_facet Li, Lingyao
Xiong, Junjie
Zhu, Changjia
Yu, Runlong
Chen, Chen
Wang, Junyu
Ma, Renkai
Lu, Zhicong
contents Large language models (LLMs) are increasingly used in academic peer review, yet their reliability, alignment with human judgment, and robustness to adversarial attacks remain poorly understood. We present a systematic benchmark of LLM-as-a-Reviewer on 898 papers stratified from NeurIPS and ICLR, evaluating 12 LLMs along three axes: rating calibration, divergence from human reviewers, and resistance to prompt injection embedded via an invisible font-mapping attack. We find that LLMs systematically overrate weaker submissions and diverge from humans in topical emphasis, under-flagging Clarity and over-flagging Reproducibility, while producing reviews two to three times longer with lower lexical diversity and a more standardized vocabulary. Prompt injection remains highly effective. Simple hidden instructions can promote low-scoring papers to acceptance-level ratings in a substantial fraction of cases, with effectiveness varying sharply across model families. While LLMs offer utility in structuring evaluations, their integration into peer review requires safeguards against both intrinsic biases and adversarial risks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_25415
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers
Li, Lingyao
Xiong, Junjie
Zhu, Changjia
Yu, Runlong
Chen, Chen
Wang, Junyu
Ma, Renkai
Lu, Zhicong
Computation and Language
Computers and Society
Emerging Technologies
Large language models (LLMs) are increasingly used in academic peer review, yet their reliability, alignment with human judgment, and robustness to adversarial attacks remain poorly understood. We present a systematic benchmark of LLM-as-a-Reviewer on 898 papers stratified from NeurIPS and ICLR, evaluating 12 LLMs along three axes: rating calibration, divergence from human reviewers, and resistance to prompt injection embedded via an invisible font-mapping attack. We find that LLMs systematically overrate weaker submissions and diverge from humans in topical emphasis, under-flagging Clarity and over-flagging Reproducibility, while producing reviews two to three times longer with lower lexical diversity and a more standardized vocabulary. Prompt injection remains highly effective. Simple hidden instructions can promote low-scoring papers to acceptance-level ratings in a substantial fraction of cases, with effectiveness varying sharply across model families. While LLMs offer utility in structuring evaluations, their integration into peer review requires safeguards against both intrinsic biases and adversarial risks.
title LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers
topic Computation and Language
Computers and Society
Emerging Technologies
url https://arxiv.org/abs/2605.25415