Certifying LLM Safety against Adversarial Prompting
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kumar, Aounon, Agarwal, Chirag, Srinivas, Suraj, Li, Aaron Jiaxun, Feizi, Soheil, Lakkaraju, Himabindu |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Detecting LLM-Generated Peer Reviews
par: Rao, Vishisht, et autres
Publié: (2025)
par: Rao, Vishisht, et autres
Publié: (2025)
MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models
par: Han, Tessa, et autres
Publié: (2024)
par: Han, Tessa, et autres
Publié: (2024)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
par: Qi, Zhenting, et autres
Publié: (2024)
par: Qi, Zhenting, et autres
Publié: (2024)
In-Context Unlearning: Language Models as Few Shot Unlearners
par: Pawelczyk, Martin, et autres
Publié: (2023)
par: Pawelczyk, Martin, et autres
Publié: (2023)
Manipulating Large Language Models to Increase Product Visibility
par: Kumar, Aounon, et autres
Publié: (2024)
par: Kumar, Aounon, et autres
Publié: (2024)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
par: Li, Aaron J., et autres
Publié: (2025)
par: Li, Aaron J., et autres
Publié: (2025)
Fast Adversarial Attacks on Language Models In One GPU Minute
par: Sadasivan, Vinu Sankar, et autres
Publié: (2024)
par: Sadasivan, Vinu Sankar, et autres
Publié: (2024)
Towards Understanding the Robustness of Sparse Autoencoders
par: Saiyed, Ahson, et autres
Publié: (2026)
par: Saiyed, Ahson, et autres
Publié: (2026)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
par: Lobo, Elita, et autres
Publié: (2024)
par: Lobo, Elita, et autres
Publié: (2024)
Tool Preferences in Agentic LLMs are Unreliable
par: Faghih, Kazem, et autres
Publié: (2025)
par: Faghih, Kazem, et autres
Publié: (2025)
Query-Based Adversarial Prompt Generation
par: Hayase, Jonathan, et autres
Publié: (2024)
par: Hayase, Jonathan, et autres
Publié: (2024)
Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry
par: Saha, Shoumik, et autres
Publié: (2026)
par: Saha, Shoumik, et autres
Publié: (2026)
An Adversarial Perspective on Machine Unlearning for AI Safety
par: Łucki, Jakub, et autres
Publié: (2024)
par: Łucki, Jakub, et autres
Publié: (2024)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
par: Li, Hao, et autres
Publié: (2026)
par: Li, Hao, et autres
Publié: (2026)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
par: Paulus, Anselm, et autres
Publié: (2024)
par: Paulus, Anselm, et autres
Publié: (2024)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
par: Mo, Yichuan, et autres
Publié: (2024)
par: Mo, Yichuan, et autres
Publié: (2024)
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification
par: Gubri, Martin, et autres
Publié: (2024)
par: Gubri, Martin, et autres
Publié: (2024)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
par: Chen, Taiye, et autres
Publié: (2025)
par: Chen, Taiye, et autres
Publié: (2025)
RECAP: A Resource-Efficient Method for Adversarial Prompting in Large Language Models
par: Chugh, Rishit
Publié: (2026)
par: Chugh, Rishit
Publié: (2026)
Prompt Leakage effect and defense strategies for multi-turn LLM interactions
par: Agarwal, Divyansh, et autres
Publié: (2024)
par: Agarwal, Divyansh, et autres
Publié: (2024)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
par: Reddy, Aashray, et autres
Publié: (2025)
par: Reddy, Aashray, et autres
Publié: (2025)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
par: Liang, Buyun, et autres
Publié: (2026)
par: Liang, Buyun, et autres
Publié: (2026)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
par: Benjamin, Victoria, et autres
Publié: (2024)
par: Benjamin, Victoria, et autres
Publié: (2024)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
par: Wu, Tong, et autres
Publié: (2024)
par: Wu, Tong, et autres
Publié: (2024)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
par: Kumar, Anurakt, et autres
Publié: (2024)
par: Kumar, Anurakt, et autres
Publié: (2024)
Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
par: Zhao, Andrew, et autres
Publié: (2025)
par: Zhao, Andrew, et autres
Publié: (2025)
RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content
par: Yuan, Zhuowen, et autres
Publié: (2024)
par: Yuan, Zhuowen, et autres
Publié: (2024)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
par: Zhang, Mohan, et autres
Publié: (2026)
par: Zhang, Mohan, et autres
Publié: (2026)
Prompt Injection attack against LLM-integrated Applications
par: Liu, Yi, et autres
Publié: (2023)
par: Liu, Yi, et autres
Publié: (2023)
In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement
par: Shetty, Anudeex, et autres
Publié: (2026)
par: Shetty, Anudeex, et autres
Publié: (2026)
Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
par: To, Bang Trinh Tran, et autres
Publié: (2025)
par: To, Bang Trinh Tran, et autres
Publié: (2025)
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
par: Chen, Sixu, et autres
Publié: (2026)
par: Chen, Sixu, et autres
Publié: (2026)
Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
par: Tang, Haochun, et autres
Publié: (2026)
par: Tang, Haochun, et autres
Publié: (2026)
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
par: Wang, Kun, et autres
Publié: (2025)
par: Wang, Kun, et autres
Publié: (2025)
Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis
par: Zhang, Haoyu, et autres
Publié: (2026)
par: Zhang, Haoyu, et autres
Publié: (2026)
LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents
par: Patel, Hitesh Laxmichand, et autres
Publié: (2024)
par: Patel, Hitesh Laxmichand, et autres
Publié: (2024)
Explaining the Model, Protecting Your Data: Revealing and Mitigating the Data Privacy Risks of Post-Hoc Model Explanations via Membership Inference
par: Huang, Catherine, et autres
Publié: (2024)
par: Huang, Catherine, et autres
Publié: (2024)
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
par: Lin, Lixing, et autres
Publié: (2026)
par: Lin, Lixing, et autres
Publié: (2026)
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
par: Jia, Feiran, et autres
Publié: (2024)
par: Jia, Feiran, et autres
Publié: (2024)
A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy
par: Correia, Pedro H. Barcha, et autres
Publié: (2026)
par: Correia, Pedro H. Barcha, et autres
Publié: (2026)
Documents similaires
-
Detecting LLM-Generated Peer Reviews
par: Rao, Vishisht, et autres
Publié: (2025) -
MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models
par: Han, Tessa, et autres
Publié: (2024) -
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
par: Qi, Zhenting, et autres
Publié: (2024) -
In-Context Unlearning: Language Models as Few Shot Unlearners
par: Pawelczyk, Martin, et autres
Publié: (2023) -
Manipulating Large Language Models to Increase Product Visibility
par: Kumar, Aounon, et autres
Publié: (2024)