Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack
Fuente:
arXiv
Saved in:
| Main Authors: | Yue, Murong, Yao, Ziyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks
by: Halloran, John T., et al.
Published: (2026)
by: Halloran, John T., et al.
Published: (2026)
Your Agent Can Defend Itself against Backdoor Attacks
by: Changjiang, Li, et al.
Published: (2025)
by: Changjiang, Li, et al.
Published: (2025)
Learning to Defend by Attacking (and Vice-Versa): Transfer of Learning in Cybersecurity Games
by: Malloy, Tailia, et al.
Published: (2023)
by: Malloy, Tailia, et al.
Published: (2023)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
by: Cao, Bochuan, et al.
Published: (2023)
by: Cao, Bochuan, et al.
Published: (2023)
Defending the Edge: Representative-Attention Defense against Backdoor Attacks in Federated Learning
by: Obioma, Chibueze Peace, et al.
Published: (2025)
by: Obioma, Chibueze Peace, et al.
Published: (2025)
Prompt-Induced Over-Generation as Denial-of-Service: A Black-Box Attack-Side Benchmark
by: Manu, et al.
Published: (2025)
by: Manu, et al.
Published: (2025)
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection
by: Chen, Shuhao, et al.
Published: (2026)
by: Chen, Shuhao, et al.
Published: (2026)
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
by: Jia, Feiran, et al.
Published: (2024)
by: Jia, Feiran, et al.
Published: (2024)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
by: Liu, Yupei, et al.
Published: (2023)
by: Liu, Yupei, et al.
Published: (2023)
Exploiting Layer-Specific Vulnerabilities to Backdoor Attack in Federated Learning
by: Foroughi, Mohammad Hadi, et al.
Published: (2026)
by: Foroughi, Mohammad Hadi, et al.
Published: (2026)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
by: Benjamin, Victoria, et al.
Published: (2024)
by: Benjamin, Victoria, et al.
Published: (2024)
LLM Security: Vulnerabilities, Attacks, Defenses, and Countermeasures
by: Aguilera-Martínez, Francisco, et al.
Published: (2025)
by: Aguilera-Martínez, Francisco, et al.
Published: (2025)
Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
by: Zhao, Andrew, et al.
Published: (2025)
by: Zhao, Andrew, et al.
Published: (2025)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
by: Qian, Cheng, et al.
Published: (2024)
by: Qian, Cheng, et al.
Published: (2024)
A Novel Solution for Zero-Day Attack Detection in IDS using Self-Attention and Jensen-Shannon Divergence in WGAN-GP
by: Mu, Ziyu, et al.
Published: (2026)
by: Mu, Ziyu, et al.
Published: (2026)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
by: Wang, Zhun, et al.
Published: (2026)
by: Wang, Zhun, et al.
Published: (2026)
Prompt Injection Attacks on Large Language Models in Oncology
by: Clusmann, Jan, et al.
Published: (2024)
by: Clusmann, Jan, et al.
Published: (2024)
Attention Tracker: Detecting Prompt Injection Attacks in LLMs
by: Hung, Kuo-Han, et al.
Published: (2024)
by: Hung, Kuo-Han, et al.
Published: (2024)
Syntax- and Compilation-Preserving Evasion of LLM Vulnerability Detectors
by: Sun, Luze, et al.
Published: (2026)
by: Sun, Luze, et al.
Published: (2026)
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
by: Li, Yuxi, et al.
Published: (2024)
by: Li, Yuxi, et al.
Published: (2024)
MoCo-EA: Exploiting Adversarial Mode Connectivity for Efficient Evolutionary Attacks
by: Kim, Hyo Seo, et al.
Published: (2026)
by: Kim, Hyo Seo, et al.
Published: (2026)
Attacks and Defenses Against LLM Fingerprinting
by: Kurian, Kevin, et al.
Published: (2025)
by: Kurian, Kevin, et al.
Published: (2025)
How Not to Detect Prompt Injections with an LLM
by: Choudhary, Sarthak, et al.
Published: (2025)
by: Choudhary, Sarthak, et al.
Published: (2025)
PLeak: Prompt Leaking Attacks against Large Language Model Applications
by: Hui, Bo, et al.
Published: (2024)
by: Hui, Bo, et al.
Published: (2024)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
by: Jiang, Weisen, et al.
Published: (2025)
by: Jiang, Weisen, et al.
Published: (2025)
Logic layer Prompt Control Injection (LPCI): A Novel Security Vulnerability Class in Agentic Systems
by: Atta, Hammad, et al.
Published: (2025)
by: Atta, Hammad, et al.
Published: (2025)
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
by: Casper, Stephen, et al.
Published: (2024)
by: Casper, Stephen, et al.
Published: (2024)
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
by: Wu, Yuanwei, et al.
Published: (2023)
by: Wu, Yuanwei, et al.
Published: (2023)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
by: Zhang, Heyi, et al.
Published: (2025)
by: Zhang, Heyi, et al.
Published: (2025)
Defending against Adversarial Malware Attacks on ML-based Android Malware Detection Systems
by: He, Ping, et al.
Published: (2025)
by: He, Ping, et al.
Published: (2025)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
by: Zhang, Mohan, et al.
Published: (2026)
by: Zhang, Mohan, et al.
Published: (2026)
LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
by: Panebianco, Francesco, et al.
Published: (2025)
by: Panebianco, Francesco, et al.
Published: (2025)
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
by: Cornacchia, Giandomenico, et al.
Published: (2024)
by: Cornacchia, Giandomenico, et al.
Published: (2024)
Attackers Strike Back? Not Anymore -- An Ensemble of RL Defenders Awakens for APT Detection
by: Benabderrahmane, Sidahmed, et al.
Published: (2025)
by: Benabderrahmane, Sidahmed, et al.
Published: (2025)
Optimal Defender Strategies for CAGE-2 using Causal Modeling and Tree Search
by: Hammar, Kim, et al.
Published: (2024)
by: Hammar, Kim, et al.
Published: (2024)
LLM Benchmark Datasets Should Be Contamination-Resistant
by: Al-Lawati, Ali, et al.
Published: (2026)
by: Al-Lawati, Ali, et al.
Published: (2026)
SurvAttack: Black-Box Attack On Survival Models through Ontology-Informed EHR Perturbation
by: Kerdabadi, Mohsen Nayebi, et al.
Published: (2024)
by: Kerdabadi, Mohsen Nayebi, et al.
Published: (2024)
Demystifying the Mythos or Disrupting Bugonomics? From Zero-Day Asymmetry to Defender Remediation Throughput
by: Pesoli, Alfredo, et al.
Published: (2026)
by: Pesoli, Alfredo, et al.
Published: (2026)
Have You Poisoned My Data? Defending Neural Networks against Data Poisoning
by: De Gaspari, Fabio, et al.
Published: (2024)
by: De Gaspari, Fabio, et al.
Published: (2024)
Fewer Weights, More Problems: A Practical Attack on LLM Pruning
by: Egashira, Kazuki, et al.
Published: (2025)
by: Egashira, Kazuki, et al.
Published: (2025)
Similar Items
-
Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks
by: Halloran, John T., et al.
Published: (2026) -
Your Agent Can Defend Itself against Backdoor Attacks
by: Changjiang, Li, et al.
Published: (2025) -
Learning to Defend by Attacking (and Vice-Versa): Transfer of Learning in Cybersecurity Games
by: Malloy, Tailia, et al.
Published: (2023) -
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
by: Cao, Bochuan, et al.
Published: (2023) -
Defending the Edge: Representative-Attention Defense against Backdoor Attacks in Federated Learning
by: Obioma, Chibueze Peace, et al.
Published: (2025)