Sparse Autoencoders are Capable LLM Jailbreak Mitigators
Fuente:
arXiv
Salvato in:
| Autori principali: | Assogba, Yannick, Cortellazzi, Jacopo, Abad, Javier, Rodriguez, Pau, Suau, Xavier, Blaas, Arno |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
di: Zeng, Yifan, et al.
Pubblicazione: (2024)
di: Zeng, Yifan, et al.
Pubblicazione: (2024)
Towards Understanding the Robustness of Sparse Autoencoders
di: Saiyed, Ahson, et al.
Pubblicazione: (2026)
di: Saiyed, Ahson, et al.
Pubblicazione: (2026)
Copyright-Protected Language Generation via Adaptive Model Fusion
di: Abad, Javier, et al.
Pubblicazione: (2024)
di: Abad, Javier, et al.
Pubblicazione: (2024)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
di: Li, Nathaniel, et al.
Pubblicazione: (2024)
di: Li, Nathaniel, et al.
Pubblicazione: (2024)
A StrongREJECT for Empty Jailbreaks
di: Souly, Alexandra, et al.
Pubblicazione: (2024)
di: Souly, Alexandra, et al.
Pubblicazione: (2024)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
di: Zhu, Sicheng, et al.
Pubblicazione: (2024)
di: Zhu, Sicheng, et al.
Pubblicazione: (2024)
Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization
di: Dotsinski, Asen, et al.
Pubblicazione: (2026)
di: Dotsinski, Asen, et al.
Pubblicazione: (2026)
Testing the Limits of Jailbreaking Defenses with the Purple Problem
di: Kim, Taeyoun, et al.
Pubblicazione: (2024)
di: Kim, Taeyoun, et al.
Pubblicazione: (2024)
Universal Jailbreak Backdoors from Poisoned Human Feedback
di: Rando, Javier, et al.
Pubblicazione: (2023)
di: Rando, Javier, et al.
Pubblicazione: (2023)
Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
di: Jia, Xiaojun, et al.
Pubblicazione: (2024)
di: Jia, Xiaojun, et al.
Pubblicazione: (2024)
VERA: Variational Inference Framework for Jailbreaking Large Language Models
di: Lochab, Anamika, et al.
Pubblicazione: (2025)
di: Lochab, Anamika, et al.
Pubblicazione: (2025)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
di: Li, Qizhang, et al.
Pubblicazione: (2024)
di: Li, Qizhang, et al.
Pubblicazione: (2024)
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
di: Hu, Hanjiang, et al.
Pubblicazione: (2025)
di: Hu, Hanjiang, et al.
Pubblicazione: (2025)
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
di: Li, Ran, et al.
Pubblicazione: (2025)
di: Li, Ran, et al.
Pubblicazione: (2025)
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
di: Kim, Heegyu, et al.
Pubblicazione: (2024)
di: Kim, Heegyu, et al.
Pubblicazione: (2024)
PromptScreen: Efficient Jailbreak Mitigation Using Semantic Linear Classification in a Multi-Staged Pipeline
di: Rao, Akshaj Prashanth, et al.
Pubblicazione: (2025)
di: Rao, Akshaj Prashanth, et al.
Pubblicazione: (2025)
From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
di: Zhang, Zhexin, et al.
Pubblicazione: (2024)
di: Zhang, Zhexin, et al.
Pubblicazione: (2024)
RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking
di: Jiang, Yifan, et al.
Pubblicazione: (2024)
di: Jiang, Yifan, et al.
Pubblicazione: (2024)
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
di: Hu, Kai, et al.
Pubblicazione: (2025)
di: Hu, Kai, et al.
Pubblicazione: (2025)
STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents
di: Li, Jing-Jing, et al.
Pubblicazione: (2025)
di: Li, Jing-Jing, et al.
Pubblicazione: (2025)
Jailbreaking in the Haystack
di: Shah, Rishi Rajesh, et al.
Pubblicazione: (2025)
di: Shah, Rishi Rajesh, et al.
Pubblicazione: (2025)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
di: Chu, Junjie, et al.
Pubblicazione: (2024)
di: Chu, Junjie, et al.
Pubblicazione: (2024)
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling
di: Ma, Avery, et al.
Pubblicazione: (2025)
di: Ma, Avery, et al.
Pubblicazione: (2025)
Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities
di: Geng, Jiahui, et al.
Pubblicazione: (2025)
di: Geng, Jiahui, et al.
Pubblicazione: (2025)
PrisonBreak: Jailbreaking Large Language Models with at Most Twenty-Five Targeted Bit-flips
di: Coalson, Zachary, et al.
Pubblicazione: (2024)
di: Coalson, Zachary, et al.
Pubblicazione: (2024)
TwinGate: Stateful Defense against Decompositional Jailbreaks in Untraceable Traffic via Asymmetric Contrastive Learning
di: Sun, Bowen, et al.
Pubblicazione: (2026)
di: Sun, Bowen, et al.
Pubblicazione: (2026)
Intriguing Properties of Adversarial ML Attacks in the Problem Space [Extended Version]
di: Cortellazzi, Jacopo, et al.
Pubblicazione: (2019)
di: Cortellazzi, Jacopo, et al.
Pubblicazione: (2019)
On the Effectiveness of Adversarial Training on Malware Classifiers
di: Bostani, Hamid, et al.
Pubblicazione: (2024)
di: Bostani, Hamid, et al.
Pubblicazione: (2024)
CodeCloak: A Method for Evaluating and Mitigating Code Leakage by LLM Code Assistants
di: Noah, Amit Finkman, et al.
Pubblicazione: (2024)
di: Noah, Amit Finkman, et al.
Pubblicazione: (2024)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
di: Rando, Javier, et al.
Pubblicazione: (2024)
di: Rando, Javier, et al.
Pubblicazione: (2024)
Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses
di: Ahmed, Mohamed, et al.
Pubblicazione: (2025)
di: Ahmed, Mohamed, et al.
Pubblicazione: (2025)
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
di: Chen, Shuo, et al.
Pubblicazione: (2024)
di: Chen, Shuo, et al.
Pubblicazione: (2024)
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization
di: Fang, Zheng, et al.
Pubblicazione: (2026)
di: Fang, Zheng, et al.
Pubblicazione: (2026)
A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy
di: Correia, Pedro H. Barcha, et al.
Pubblicazione: (2026)
di: Correia, Pedro H. Barcha, et al.
Pubblicazione: (2026)
TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice
di: Goel, Aman, et al.
Pubblicazione: (2025)
di: Goel, Aman, et al.
Pubblicazione: (2025)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
di: Panfilov, Alexander, et al.
Pubblicazione: (2025)
di: Panfilov, Alexander, et al.
Pubblicazione: (2025)
Jailbreaking LLMs via Calibration
di: Lu, Yuxuan, et al.
Pubblicazione: (2026)
di: Lu, Yuxuan, et al.
Pubblicazione: (2026)
Jailbreaking with Universal Multi-Prompts
di: Hsu, Yu-Ling, et al.
Pubblicazione: (2025)
di: Hsu, Yu-Ling, et al.
Pubblicazione: (2025)
ContinuousBench: Can Differentially Private Synthetic Text Improve Capabilities?
di: Liu, Peihan, et al.
Pubblicazione: (2026)
di: Liu, Peihan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025) -
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
di: Zeng, Yifan, et al.
Pubblicazione: (2024) -
Towards Understanding the Robustness of Sparse Autoencoders
di: Saiyed, Ahson, et al.
Pubblicazione: (2026) -
Copyright-Protected Language Generation via Adaptive Model Fusion
di: Abad, Javier, et al.
Pubblicazione: (2024) -
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
di: Li, Nathaniel, et al.
Pubblicazione: (2024)