Breaking the Reviewer: Assessing the Vulnerability of Large Language Models in Automated Peer Review Under Textual Adversarial Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Tzu-Ling, Chen, Wei-Chih, Hsiao, Teng-Fang, Liu, Hou-I, Yeh, Ya-Hsin, Chan, Yu Kai, Lien, Wen-Sheng, Kuo, Po-Yen, Yu, Philip S., Shuai, Hong-Han
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909832000307200
author Lin, Tzu-Ling
Chen, Wei-Chih
Hsiao, Teng-Fang
Liu, Hou-I
Yeh, Ya-Hsin
Chan, Yu Kai
Lien, Wen-Sheng
Kuo, Po-Yen
Yu, Philip S.
Shuai, Hong-Han
author_facet Lin, Tzu-Ling
Chen, Wei-Chih
Hsiao, Teng-Fang
Liu, Hou-I
Yeh, Ya-Hsin
Chan, Yu Kai
Lien, Wen-Sheng
Kuo, Po-Yen
Yu, Philip S.
Shuai, Hong-Han
contents Peer review is essential for maintaining academic quality, but the increasing volume of submissions places a significant burden on reviewers. Large language models (LLMs) offer potential assistance in this process, yet their susceptibility to textual adversarial attacks raises reliability concerns. This paper investigates the robustness of LLMs used as automated reviewers in the presence of such attacks. We focus on three key questions: (1) The effectiveness of LLMs in generating reviews compared to human reviewers. (2) The impact of adversarial attacks on the reliability of LLM-generated reviews. (3) Challenges and potential mitigation strategies for LLM-based review. Our evaluation reveals significant vulnerabilities, as text manipulations can distort LLM assessments. We offer a comprehensive evaluation of LLM performance in automated peer reviewing and analyze its robustness against adversarial attacks. Our findings emphasize the importance of addressing adversarial risks to ensure AI strengthens, rather than compromises, the integrity of scholarly communication.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11113
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Breaking the Reviewer: Assessing the Vulnerability of Large Language Models in Automated Peer Review Under Textual Adversarial Attacks
Lin, Tzu-Ling
Chen, Wei-Chih
Hsiao, Teng-Fang
Liu, Hou-I
Yeh, Ya-Hsin
Chan, Yu Kai
Lien, Wen-Sheng
Kuo, Po-Yen
Yu, Philip S.
Shuai, Hong-Han
Computation and Language
Artificial Intelligence
Peer review is essential for maintaining academic quality, but the increasing volume of submissions places a significant burden on reviewers. Large language models (LLMs) offer potential assistance in this process, yet their susceptibility to textual adversarial attacks raises reliability concerns. This paper investigates the robustness of LLMs used as automated reviewers in the presence of such attacks. We focus on three key questions: (1) The effectiveness of LLMs in generating reviews compared to human reviewers. (2) The impact of adversarial attacks on the reliability of LLM-generated reviews. (3) Challenges and potential mitigation strategies for LLM-based review. Our evaluation reveals significant vulnerabilities, as text manipulations can distort LLM assessments. We offer a comprehensive evaluation of LLM performance in automated peer reviewing and analyze its robustness against adversarial attacks. Our findings emphasize the importance of addressing adversarial risks to ensure AI strengthens, rather than compromises, the integrity of scholarly communication.
title Breaking the Reviewer: Assessing the Vulnerability of Large Language Models in Automated Peer Review Under Textual Adversarial Attacks
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.11113