PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
Fuente:
arXiv
Salvato in:
| Autori principali: | Mangaokar, Neal, Hooda, Ashish, Choi, Jihye, Chandrashekaran, Shreyas, Fawaz, Kassem, Jha, Somesh, Prakash, Atul |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Defending Object Detectors against Patch Attacks with Out-of-Distribution Smoothing
di: Feng, Ryan, et al.
Pubblicazione: (2022)
di: Feng, Ryan, et al.
Pubblicazione: (2022)
What Really is a Member? Discrediting Membership Inference via Poisoning
di: Mangaokar, Neal, et al.
Pubblicazione: (2025)
di: Mangaokar, Neal, et al.
Pubblicazione: (2025)
PolicyLR: A Logic Representation For Privacy Policies
di: Hooda, Ashish, et al.
Pubblicazione: (2024)
di: Hooda, Ashish, et al.
Pubblicazione: (2024)
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
di: Choudhary, Sarthak, et al.
Pubblicazione: (2026)
di: Choudhary, Sarthak, et al.
Pubblicazione: (2026)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
di: Wang, Zi, et al.
Pubblicazione: (2024)
di: Wang, Zi, et al.
Pubblicazione: (2024)
Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025)
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025)
SEA: Shareable and Explainable Attribution for Query-based Black-box Attacks
di: Gao, Yue, et al.
Pubblicazione: (2023)
di: Gao, Yue, et al.
Pubblicazione: (2023)
SLVR: Securely Leveraging Client Validation for Robust Federated Learning
di: Choi, Jihye, et al.
Pubblicazione: (2025)
di: Choi, Jihye, et al.
Pubblicazione: (2025)
Systems Security Foundations for Agentic Computing
di: Christodorescu, Mihai, et al.
Pubblicazione: (2025)
di: Christodorescu, Mihai, et al.
Pubblicazione: (2025)
Taming Data Challenges in ML-based Security Tasks Using Generative AI
di: Kanchi, Shravya, et al.
Pubblicazione: (2025)
di: Kanchi, Shravya, et al.
Pubblicazione: (2025)
Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and Conditions in Shopping Websites at Scale
di: Tsai, Elisa, et al.
Pubblicazione: (2025)
di: Tsai, Elisa, et al.
Pubblicazione: (2025)
Agent Security is a Systems Problem
di: Christodorescu, Mihai, et al.
Pubblicazione: (2026)
di: Christodorescu, Mihai, et al.
Pubblicazione: (2026)
Formal Policy Enforcement for Real-World Agentic Systems
di: Palumbo, Nils, et al.
Pubblicazione: (2026)
di: Palumbo, Nils, et al.
Pubblicazione: (2026)
Prediction with Expert Advice under Local Differential Privacy
di: Jacobsen, Ben, et al.
Pubblicazione: (2025)
di: Jacobsen, Ben, et al.
Pubblicazione: (2025)
Private Continual Counting of Unbounded Streams
di: Jacobsen, Ben, et al.
Pubblicazione: (2025)
di: Jacobsen, Ben, et al.
Pubblicazione: (2025)
Confused ChatGPT: Cross-App Context Poisoning via First-Party APIs
di: Wang, Chao, et al.
Pubblicazione: (2026)
di: Wang, Chao, et al.
Pubblicazione: (2026)
ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning
di: Zhao, Zhengyue, et al.
Pubblicazione: (2025)
di: Zhao, Zhengyue, et al.
Pubblicazione: (2025)
CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2025)
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2025)
Fun-tuning: Characterizing the Vulnerability of Proprietary LLMs to Optimization-based Prompt Injection Attacks via the Fine-Tuning Interface
di: Labunets, Andrey, et al.
Pubblicazione: (2025)
di: Labunets, Andrey, et al.
Pubblicazione: (2025)
Dependency-Aware Privacy for Multi-turn Agents
di: Anshumaan, Divyam, et al.
Pubblicazione: (2026)
di: Anshumaan, Divyam, et al.
Pubblicazione: (2026)
JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks
di: Zhang, Xiaoyu, et al.
Pubblicazione: (2023)
di: Zhang, Xiaoyu, et al.
Pubblicazione: (2023)
Personalizing Agent Privacy Decisions via Logical Entailment
di: Flemings, James, et al.
Pubblicazione: (2025)
di: Flemings, James, et al.
Pubblicazione: (2025)
SmoothGuard: Defending Multimodal Large Language Models with Noise Perturbation and Clustering Aggregation
di: Su, Guangzhi, et al.
Pubblicazione: (2025)
di: Su, Guangzhi, et al.
Pubblicazione: (2025)
SmartGuard: Leveraging Large Language Models for Network Attack Detection through Audit Log Analysis and Summarization
di: Zhang, Hao, et al.
Pubblicazione: (2025)
di: Zhang, Hao, et al.
Pubblicazione: (2025)
Text-Based Personas for Simulating User Privacy Decisions
di: Fawaz, Kassem, et al.
Pubblicazione: (2026)
di: Fawaz, Kassem, et al.
Pubblicazione: (2026)
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
di: Chen, Yulin, et al.
Pubblicazione: (2026)
di: Chen, Yulin, et al.
Pubblicazione: (2026)
Using Retriever Augmented Large Language Models for Attack Graph Generation
di: Prapty, Renascence Tarafder, et al.
Pubblicazione: (2024)
di: Prapty, Renascence Tarafder, et al.
Pubblicazione: (2024)
WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks
di: Ramesh, Guruprasad Viswanathan, et al.
Pubblicazione: (2026)
di: Ramesh, Guruprasad Viswanathan, et al.
Pubblicazione: (2026)
DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
di: Li, Hao, et al.
Pubblicazione: (2025)
di: Li, Hao, et al.
Pubblicazione: (2025)
Can One Safety Loop Guard Them All? Agentic Guard Rails for Federated Computing
di: Veeraragavan, Narasimha Raghavan, et al.
Pubblicazione: (2025)
di: Veeraragavan, Narasimha Raghavan, et al.
Pubblicazione: (2025)
TraceGuard: Process-Guided Firewall against Reasoning Backdoors in Large Language Models
di: Guo, Zhen, et al.
Pubblicazione: (2026)
di: Guo, Zhen, et al.
Pubblicazione: (2026)
Do Large Code Models Understand Programming Concepts? Counterfactual Analysis for Code Predicates
di: Hooda, Ashish, et al.
Pubblicazione: (2024)
di: Hooda, Ashish, et al.
Pubblicazione: (2024)
SAGE: Sample-Aware Guarding Engine for Robust Intrusion Detection Against Adversarial Attacks
di: Chen, Jing, et al.
Pubblicazione: (2025)
di: Chen, Jing, et al.
Pubblicazione: (2025)
KinGuard: Hierarchical Kinship-Aware Fingerprinting to Defend Against Large Language Model Stealing
di: Xu, Zhenhua, et al.
Pubblicazione: (2026)
di: Xu, Zhenhua, et al.
Pubblicazione: (2026)
ProP: Efficient Backdoor Detection via Propagation Perturbation for Overparametrized Models
di: Ren, Tao, et al.
Pubblicazione: (2024)
di: Ren, Tao, et al.
Pubblicazione: (2024)
Publicly-Detectable Watermarking for Language Models
di: Fairoze, Jaiden, et al.
Pubblicazione: (2023)
di: Fairoze, Jaiden, et al.
Pubblicazione: (2023)
How Not to Detect Prompt Injections with an LLM
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025)
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025)
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
di: Wu, Fangzhou, et al.
Pubblicazione: (2024)
di: Wu, Fangzhou, et al.
Pubblicazione: (2024)
On the Difficulty of Constructing a Robust and Publicly-Detectable Watermark
di: Fairoze, Jaiden, et al.
Pubblicazione: (2025)
di: Fairoze, Jaiden, et al.
Pubblicazione: (2025)
Less Is More: Sparse and Cooperative Perturbation for Point Cloud Attacks
di: Tang, Keke, et al.
Pubblicazione: (2025)
di: Tang, Keke, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Defending Object Detectors against Patch Attacks with Out-of-Distribution Smoothing
di: Feng, Ryan, et al.
Pubblicazione: (2022) -
What Really is a Member? Discrediting Membership Inference via Poisoning
di: Mangaokar, Neal, et al.
Pubblicazione: (2025) -
PolicyLR: A Logic Representation For Privacy Policies
di: Hooda, Ashish, et al.
Pubblicazione: (2024) -
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
di: Choudhary, Sarthak, et al.
Pubblicazione: (2026) -
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
di: Wang, Zi, et al.
Pubblicazione: (2024)