Multilingual Hidden Prompt Injection Attacks on LLM-Based Academic Reviewing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Theocharopoulos, Panagiotis, Kulkarni, Ajinkya, -Doss, Mathew Magimai.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909977856180224
author Theocharopoulos, Panagiotis
Kulkarni, Ajinkya
-Doss, Mathew Magimai.
author_facet Theocharopoulos, Panagiotis
Kulkarni, Ajinkya
-Doss, Mathew Magimai.
contents Large language models (LLMs) are increasingly considered for use in high-impact workflows, including academic peer review. However, LLMs are vulnerable to document-level hidden prompt injection attacks. In this work, we construct a dataset of approximately 500 real academic papers accepted to ICML and evaluate the effect of embedding hidden adversarial prompts within these documents. Each paper is injected with semantically equivalent instructions in four different languages and reviewed using an LLM. We find that prompt injection induces substantial changes in review scores and accept/reject decisions for English, Japanese, and Chinese injections, while Arabic injections produce little to no effect. These results highlight the susceptibility of LLM-based reviewing systems to document-level prompt injection and reveal notable differences in vulnerability across languages.
format Preprint
id arxiv_https___arxiv_org_abs_2512_23684
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multilingual Hidden Prompt Injection Attacks on LLM-Based Academic Reviewing
Theocharopoulos, Panagiotis
Kulkarni, Ajinkya
-Doss, Mathew Magimai.
Computation and Language
Artificial Intelligence
Large language models (LLMs) are increasingly considered for use in high-impact workflows, including academic peer review. However, LLMs are vulnerable to document-level hidden prompt injection attacks. In this work, we construct a dataset of approximately 500 real academic papers accepted to ICML and evaluate the effect of embedding hidden adversarial prompts within these documents. Each paper is injected with semantically equivalent instructions in four different languages and reviewed using an LLM. We find that prompt injection induces substantial changes in review scores and accept/reject decisions for English, Japanese, and Chinese injections, while Arabic injections produce little to no effect. These results highlight the susceptibility of LLM-based reviewing systems to document-level prompt injection and reveal notable differences in vulnerability across languages.
title Multilingual Hidden Prompt Injection Attacks on LLM-Based Academic Reviewing
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.23684