SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xin, Yuan, Weng, Yixuan, Zhu, Minjun, Ling, Ying, Qin, Chengwei, Backes, Michael, Zhang, Yue, Yang, Linyi
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918528701956096
author Xin, Yuan
Weng, Yixuan
Zhu, Minjun
Ling, Ying
Qin, Chengwei
Backes, Michael
Zhang, Yue
Yang, Linyi
author_facet Xin, Yuan
Weng, Yixuan
Zhu, Minjun
Ling, Ying
Qin, Chengwei
Backes, Michael
Zhang, Yue
Yang, Linyi
contents As Large Language Models (LLMs) are increasingly integrated into academic peer review, their vulnerability to adversarial hidden prompts, i.e., adversarial instructions embedded in submissions to manipulate outcomes, poses a critical threat to scholarly integrity. We propose SafeReview, a co-evolutionary adversarial training framework for defending LLM-based peer review systems against such attacks. SafeReview jointly trains a Generator model to create sophisticated attack prompts and a Defender model to preserve review integrity under adversarial manipulation. The Generator is optimized to produce increasingly effective prompt injections, while the Defender is strengthened through preference-based training to maintain consistent reviews between clean and attacked submissions. Experimental results show that SafeReview improves robustness against adaptive prompt injection attacks, better preserves paper ranking under attack, and generalizes across attacker architectures compared with static defenses. These results demonstrate the potential of co-evolutionary training as a foundation for securing LLM-assisted peer review.
format Preprint
id arxiv_https___arxiv_org_abs_2604_26506
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
Xin, Yuan
Weng, Yixuan
Zhu, Minjun
Ling, Ying
Qin, Chengwei
Backes, Michael
Zhang, Yue
Yang, Linyi
Computation and Language
Cryptography and Security
As Large Language Models (LLMs) are increasingly integrated into academic peer review, their vulnerability to adversarial hidden prompts, i.e., adversarial instructions embedded in submissions to manipulate outcomes, poses a critical threat to scholarly integrity. We propose SafeReview, a co-evolutionary adversarial training framework for defending LLM-based peer review systems against such attacks. SafeReview jointly trains a Generator model to create sophisticated attack prompts and a Defender model to preserve review integrity under adversarial manipulation. The Generator is optimized to produce increasingly effective prompt injections, while the Defender is strengthened through preference-based training to maintain consistent reviews between clean and attacked submissions. Experimental results show that SafeReview improves robustness against adaptive prompt injection attacks, better preserves paper ranking under attack, and generalizes across attacker architectures compared with static defenses. These results demonstrate the potential of co-evolutionary training as a foundation for securing LLM-assisted peer review.
title SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
topic Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2604.26506