Adaptive Probe-based Steering for Robust LLM Jailbreaking
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Junxi, Dong, Junhao, Xie, Xiaohua |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exploring Adversarial Attacks against Latent Diffusion Model from the Perspective of Adversarial Transferability
por: Chen, Junxi, et al.
Publicado: (2024)
por: Chen, Junxi, et al.
Publicado: (2024)
Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive Jailbreaking
por: Chen, Junxi, et al.
Publicado: (2025)
por: Chen, Junxi, et al.
Publicado: (2025)
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
por: Hu, Hanjiang, et al.
Publicado: (2025)
por: Hu, Hanjiang, et al.
Publicado: (2025)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
por: Chao, Patrick, et al.
Publicado: (2024)
por: Chao, Patrick, et al.
Publicado: (2024)
TrojanPraise: Jailbreak LLMs via Benign Fine-Tuning
por: Xie, Zhixin, et al.
Publicado: (2026)
por: Xie, Zhixin, et al.
Publicado: (2026)
TRYLOCK: Defense-in-Depth Against LLM Jailbreaks via Layered Preference and Representation Engineering
por: Thornton, Scott
Publicado: (2026)
por: Thornton, Scott
Publicado: (2026)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
por: Hossain, Ismail, et al.
Publicado: (2026)
por: Hossain, Ismail, et al.
Publicado: (2026)
CARE: Ensemble Adversarial Robustness Evaluation Against Adaptive Attackers for Security Applications
por: Zhang, Hangsheng, et al.
Publicado: (2024)
por: Zhang, Hangsheng, et al.
Publicado: (2024)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
por: Nasr, Milad, et al.
Publicado: (2025)
por: Nasr, Milad, et al.
Publicado: (2025)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
por: Assogba, Yannick, et al.
Publicado: (2026)
por: Assogba, Yannick, et al.
Publicado: (2026)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
por: Li, Nathaniel, et al.
Publicado: (2024)
por: Li, Nathaniel, et al.
Publicado: (2024)
TokenProber: Jailbreaking Text-to-image Models via Fine-grained Word Impact Analysis
por: Wang, Longtian, et al.
Publicado: (2025)
por: Wang, Longtian, et al.
Publicado: (2025)
Understanding and Enhancing the Transferability of Jailbreaking Attacks
por: Lin, Runqi, et al.
Publicado: (2025)
por: Lin, Runqi, et al.
Publicado: (2025)
Adaptive Meta-learning-based Adversarial Training for Robust Automatic Modulation Classification
por: Bamdad, Amirmohammad, et al.
Publicado: (2025)
por: Bamdad, Amirmohammad, et al.
Publicado: (2025)
Security-by-Design for LLM-Based Code Generation: Leveraging Internal Representations for Concept-Driven Steering Mechanisms
por: Wendlinger, Maximilian, et al.
Publicado: (2026)
por: Wendlinger, Maximilian, et al.
Publicado: (2026)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
Jailbreak Attack Initializations as Extractors of Compliance Directions
por: Levi, Amit, et al.
Publicado: (2025)
por: Levi, Amit, et al.
Publicado: (2025)
ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs
por: Liu, Xu, et al.
Publicado: (2025)
por: Liu, Xu, et al.
Publicado: (2025)
Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models
por: Dong, Yingkai, et al.
Publicado: (2024)
por: Dong, Yingkai, et al.
Publicado: (2024)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
por: Wang, Zi, et al.
Publicado: (2024)
por: Wang, Zi, et al.
Publicado: (2024)
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
por: Shen, Xinyue, et al.
Publicado: (2023)
por: Shen, Xinyue, et al.
Publicado: (2023)
Voice Jailbreak Attacks Against GPT-4o
por: Shen, Xinyue, et al.
Publicado: (2024)
por: Shen, Xinyue, et al.
Publicado: (2024)
Jailbreaking Large Language Models in Infinitely Many Ways
por: Goldstein, Oliver, et al.
Publicado: (2025)
por: Goldstein, Oliver, et al.
Publicado: (2025)
DeepInception: Hypnotize Large Language Model to Be Jailbreaker
por: Li, Xuan, et al.
Publicado: (2023)
por: Li, Xuan, et al.
Publicado: (2023)
JULI: Jailbreak Large Language Models by Self-Introspection
por: Wang, Jesson, et al.
Publicado: (2025)
por: Wang, Jesson, et al.
Publicado: (2025)
Defending Jailbreak Prompts via In-Context Adversarial Game
por: Zhou, Yujun, et al.
Publicado: (2024)
por: Zhou, Yujun, et al.
Publicado: (2024)
Is The Watermarking Of LLM-Generated Code Robust?
por: Suresh, Tarun, et al.
Publicado: (2024)
por: Suresh, Tarun, et al.
Publicado: (2024)
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
por: Lin, Shuyi, et al.
Publicado: (2025)
por: Lin, Shuyi, et al.
Publicado: (2025)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
por: Zeng, Yifan, et al.
Publicado: (2024)
por: Zeng, Yifan, et al.
Publicado: (2024)
Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models
por: Wang, Xiangwen, et al.
Publicado: (2026)
por: Wang, Xiangwen, et al.
Publicado: (2026)
Knowledge-Driven Multi-Turn Jailbreaking on Large Language Models
por: Li, Songze, et al.
Publicado: (2026)
por: Li, Songze, et al.
Publicado: (2026)
LLMStinger: Jailbreaking LLMs using RL fine-tuned LLMs
por: Jha, Piyush, et al.
Publicado: (2024)
por: Jha, Piyush, et al.
Publicado: (2024)
Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation
por: Chen, Luoyu, et al.
Publicado: (2026)
por: Chen, Luoyu, et al.
Publicado: (2026)
Survey on Adversarial Attack and Defense for Medical Image Analysis: Methods and Challenges
por: Dong, Junhao, et al.
Publicado: (2023)
por: Dong, Junhao, et al.
Publicado: (2023)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
por: Nikolić, Kristina, et al.
Publicado: (2025)
por: Nikolić, Kristina, et al.
Publicado: (2025)
T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video Models
por: Liang, Siyuan, et al.
Publicado: (2025)
por: Liang, Siyuan, et al.
Publicado: (2025)
MetaCipher: A Time-Persistent and Universal Multi-Agent Framework for Cipher-Based Jailbreak Attacks for LLMs
por: Chen, Boyuan, et al.
Publicado: (2025)
por: Chen, Boyuan, et al.
Publicado: (2025)
LMEraser: Large Model Unlearning through Adaptive Prompt Tuning
por: Xu, Jie, et al.
Publicado: (2024)
por: Xu, Jie, et al.
Publicado: (2024)
Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual Steganography
por: Li, Songze, et al.
Publicado: (2025)
por: Li, Songze, et al.
Publicado: (2025)
SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision-Language Models
por: Zeng, Xiyu, et al.
Publicado: (2025)
por: Zeng, Xiyu, et al.
Publicado: (2025)
Ejemplares similares
-
Exploring Adversarial Attacks against Latent Diffusion Model from the Perspective of Adversarial Transferability
por: Chen, Junxi, et al.
Publicado: (2024) -
Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive Jailbreaking
por: Chen, Junxi, et al.
Publicado: (2025) -
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
por: Hu, Hanjiang, et al.
Publicado: (2025) -
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
por: Chao, Patrick, et al.
Publicado: (2024) -
TrojanPraise: Jailbreak LLMs via Benign Fine-Tuning
por: Xie, Zhixin, et al.
Publicado: (2026)