Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yue, Murong, Yao, Ziyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918065194663936
author Yue, Murong
Yao, Ziyu
author_facet Yue, Murong
Yao, Ziyu
contents Batch prompting, which combines a batch of multiple queries sharing the same context in one inference, has emerged as a promising solution to reduce inference costs. However, our study reveals a significant security vulnerability in batch prompting: malicious users can inject attack instructions into a batch, leading to unwanted interference across all queries, which can result in the inclusion of harmful content, such as phishing links, or the disruption of logical reasoning. In this paper, we construct BATCHSAFEBENCH, a comprehensive benchmark comprising 150 attack instructions of two types and 8k batch instances, to study the batch prompting vulnerability systematically. Our evaluation of both closed-source and open-weight LLMs demonstrates that all LLMs are susceptible to batch-prompting attacks. We then explore multiple defending approaches. While the prompting-based defense shows limited effectiveness for smaller LLMs, the probing-based approach achieves about 95% accuracy in detecting attacks. Additionally, we perform a mechanistic analysis to understand the attack and identify attention heads that are responsible for it.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15551
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack
Yue, Murong
Yao, Ziyu
Cryptography and Security
Artificial Intelligence
Machine Learning
Batch prompting, which combines a batch of multiple queries sharing the same context in one inference, has emerged as a promising solution to reduce inference costs. However, our study reveals a significant security vulnerability in batch prompting: malicious users can inject attack instructions into a batch, leading to unwanted interference across all queries, which can result in the inclusion of harmful content, such as phishing links, or the disruption of logical reasoning. In this paper, we construct BATCHSAFEBENCH, a comprehensive benchmark comprising 150 attack instructions of two types and 8k batch instances, to study the batch prompting vulnerability systematically. Our evaluation of both closed-source and open-weight LLMs demonstrates that all LLMs are susceptible to batch-prompting attacks. We then explore multiple defending approaches. While the prompting-based defense shows limited effectiveness for smaller LLMs, the probing-based approach achieves about 95% accuracy in detecting attacks. Additionally, we perform a mechanistic analysis to understand the attack and identify attention heads that are responsible for it.
title Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.15551