RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
Fuente:
arXiv
Saved in:
| Main Authors: | Wen, Yuxin, Zharmagambetov, Arman, Evtimov, Ivan, Kokhlikyan, Narine, Goldstein, Tom, Chaudhuri, Kamalika, Guo, Chuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
by: Evtimov, Ivan, et al.
Published: (2025)
by: Evtimov, Ivan, et al.
Published: (2025)
CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
by: Mireshghallah, Niloofar, et al.
Published: (2025)
by: Mireshghallah, Niloofar, et al.
Published: (2025)
SecAlign: Defending Against Prompt Injection with Preference Optimization
by: Chen, Sizhe, et al.
Published: (2024)
by: Chen, Sizhe, et al.
Published: (2024)
Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
by: Chen, Sizhe, et al.
Published: (2025)
by: Chen, Sizhe, et al.
Published: (2025)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
by: Paulus, Anselm, et al.
Published: (2024)
by: Paulus, Anselm, et al.
Published: (2024)
How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
by: Dziemian, Mateusz, et al.
Published: (2026)
by: Dziemian, Mateusz, et al.
Published: (2026)
Measuring Privacy Loss in Distributed Spatio-Temporal Data
by: Koga, Tatsuki, et al.
Published: (2024)
by: Koga, Tatsuki, et al.
Published: (2024)
Machine Learning with Privacy for Protected Attributes
by: Mahloujifar, Saeed, et al.
Published: (2025)
by: Mahloujifar, Saeed, et al.
Published: (2025)
Robustness of Locally Differentially Private Graph Analysis Against Poisoning
by: Imola, Jacob, et al.
Published: (2022)
by: Imola, Jacob, et al.
Published: (2022)
Metric Differential Privacy at the User-Level Via the Earth Mover's Distance
by: Imola, Jacob, et al.
Published: (2024)
by: Imola, Jacob, et al.
Published: (2024)
Communication-Efficient Triangle Counting under Local Differential Privacy
by: Imola, Jacob, et al.
Published: (2021)
by: Imola, Jacob, et al.
Published: (2021)
Privacy Amplification for the Gaussian Mechanism via Bounded Support
by: Hu, Shengyuan, et al.
Published: (2024)
by: Hu, Shengyuan, et al.
Published: (2024)
Auditing $f$-Differential Privacy in One Run
by: Mahloujifar, Saeed, et al.
Published: (2024)
by: Mahloujifar, Saeed, et al.
Published: (2024)
Can We Infer Confidential Properties of Training Data from LLMs?
by: Huang, Pengrun, et al.
Published: (2025)
by: Huang, Pengrun, et al.
Published: (2025)
Z0-Inf: Zeroth Order Approximation for Data Influence
by: Kokhlikyan, Narine, et al.
Published: (2025)
by: Kokhlikyan, Narine, et al.
Published: (2025)
PromptArmor: Simple yet Effective Prompt Injection Defenses
by: Shi, Tianneng, et al.
Published: (2025)
by: Shi, Tianneng, et al.
Published: (2025)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
by: Wang, Jiawen, et al.
Published: (2025)
by: Wang, Jiawen, et al.
Published: (2025)
Privacy Blur: Quantifying Privacy and Utility for Image Data Release
by: Mahloujifar, Saeed, et al.
Published: (2025)
by: Mahloujifar, Saeed, et al.
Published: (2025)
Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents
by: Zharmagambetov, Arman, et al.
Published: (2025)
by: Zharmagambetov, Arman, et al.
Published: (2025)
$ρ$Hammer: Reviving RowHammer Attacks on New Architectures via Prefetching
by: Chen, Weijie, et al.
Published: (2025)
by: Chen, Weijie, et al.
Published: (2025)
Fingerprinting LLMs via Prompt Injection
by: Hu, Yuepeng, et al.
Published: (2025)
by: Hu, Yuepeng, et al.
Published: (2025)
Coercing LLMs to do and reveal (almost) anything
by: Geiping, Jonas, et al.
Published: (2024)
by: Geiping, Jonas, et al.
Published: (2024)
GbHammer: Malicious Inter-process Page Sharing by Hammering Global Bits in Page Table Entries
by: Yoshioka, Keigo, et al.
Published: (2024)
by: Yoshioka, Keigo, et al.
Published: (2024)
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
by: Zhong, Yinan, et al.
Published: (2025)
by: Zhong, Yinan, et al.
Published: (2025)
Guarantees of confidentiality via Hammersley-Chapman-Robbins bounds
by: Chaudhuri, Kamalika, et al.
Published: (2024)
by: Chaudhuri, Kamalika, et al.
Published: (2024)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
by: Jaiswal, Piyush, et al.
Published: (2026)
by: Jaiswal, Piyush, et al.
Published: (2026)
Persistent Pre-Training Poisoning of LLMs
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
PromptShield: Deployable Detection for Prompt Injection Attacks
by: Jacob, Dennis, et al.
Published: (2025)
by: Jacob, Dennis, et al.
Published: (2025)
VortexPIA: Indirect Prompt Injection Attack against LLMs for Efficient Extraction of User Privacy
by: Cui, Yu, et al.
Published: (2025)
by: Cui, Yu, et al.
Published: (2025)
HammerSim: A System-Level Tool to Model RowHammer
by: Goswami, Kaustav, et al.
Published: (2026)
by: Goswami, Kaustav, et al.
Published: (2026)
Has My System Prompt Been Used? Large Language Model Prompt Membership Inference
by: Levin, Roman, et al.
Published: (2025)
by: Levin, Roman, et al.
Published: (2025)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
by: Zhu, Sicheng, et al.
Published: (2024)
by: Zhu, Sicheng, et al.
Published: (2024)
Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs
by: Yeo, Andrew, et al.
Published: (2025)
by: Yeo, Andrew, et al.
Published: (2025)
Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
by: Wen, Yuxin, et al.
Published: (2024)
by: Wen, Yuxin, et al.
Published: (2024)
Privacy-Preserving Prompt Injection Detection for LLMs Using Federated Learning and Embedding-Based NLP Classification
by: Jayathilaka, Hasini
Published: (2025)
by: Jayathilaka, Hasini
Published: (2025)
BreakHammer: Enhancing RowHammer Mitigations by Carefully Throttling Suspect Threads
by: Canpolat, Oğuzhan, et al.
Published: (2024)
by: Canpolat, Oğuzhan, et al.
Published: (2024)
Privacy-Preserving Retrieval-Augmented Generation with Differential Privacy
by: Koga, Tatsuki, et al.
Published: (2024)
by: Koga, Tatsuki, et al.
Published: (2024)
RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMs
by: Chen, Xuan, et al.
Published: (2024)
by: Chen, Xuan, et al.
Published: (2024)
Gradient-based Jailbreak Images for Multimodal Fusion Models
by: Rando, Javier, et al.
Published: (2024)
by: Rando, Javier, et al.
Published: (2024)
Similar Items
-
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
by: Evtimov, Ivan, et al.
Published: (2025) -
CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
by: Mireshghallah, Niloofar, et al.
Published: (2025) -
SecAlign: Defending Against Prompt Injection with Preference Optimization
by: Chen, Sizhe, et al.
Published: (2024) -
Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
by: Chen, Sizhe, et al.
Published: (2025) -
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
by: Paulus, Anselm, et al.
Published: (2024)