Jailbreak Attack Initializations as Extractors of Compliance Directions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Levi, Amit, Himelstein, Rom, Nemcovsky, Yaniv, Mendelson, Avi, Baskin, Chaim |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sparse patches adversarial attacks via extrapolating point-wise information
von: Nemcovsky, Yaniv, et al.
Veröffentlicht: (2024)
von: Nemcovsky, Yaniv, et al.
Veröffentlicht: (2024)
Silenced Biases: The Dark Side LLMs Learned to Refuse
von: Himelstein, Rom, et al.
Veröffentlicht: (2025)
von: Himelstein, Rom, et al.
Veröffentlicht: (2025)
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models
von: Rahimi, Eliron, et al.
Veröffentlicht: (2026)
von: Rahimi, Eliron, et al.
Veröffentlicht: (2026)
Model Inversion Attacks Meet Cryptographic Fuzzy Extractors
von: Prabhakar, Mallika, et al.
Veröffentlicht: (2025)
von: Prabhakar, Mallika, et al.
Veröffentlicht: (2025)
Indiscriminate Data Poisoning Attacks on Pre-trained Feature Extractors
von: Lu, Yiwei, et al.
Veröffentlicht: (2024)
von: Lu, Yiwei, et al.
Veröffentlicht: (2024)
Silent Tokens, Loud Effects: Padding in LLMs
von: Himelstein, Rom, et al.
Veröffentlicht: (2025)
von: Himelstein, Rom, et al.
Veröffentlicht: (2025)
Understanding and Enhancing the Transferability of Jailbreaking Attacks
von: Lin, Runqi, et al.
Veröffentlicht: (2025)
von: Lin, Runqi, et al.
Veröffentlicht: (2025)
Voice Jailbreak Attacks Against GPT-4o
von: Shen, Xinyue, et al.
Veröffentlicht: (2024)
von: Shen, Xinyue, et al.
Veröffentlicht: (2024)
Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models
von: Wang, Xiangwen, et al.
Veröffentlicht: (2026)
von: Wang, Xiangwen, et al.
Veröffentlicht: (2026)
SCART: Simulation of Cyber Attacks for Real-Time
von: Rahimi, Eliron, et al.
Veröffentlicht: (2023)
von: Rahimi, Eliron, et al.
Veröffentlicht: (2023)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
Transpose Attack: Stealing Datasets with Bidirectional Training
von: Amit, Guy, et al.
Veröffentlicht: (2023)
von: Amit, Guy, et al.
Veröffentlicht: (2023)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
von: Nasr, Milad, et al.
Veröffentlicht: (2025)
von: Nasr, Milad, et al.
Veröffentlicht: (2025)
MetaCipher: A Time-Persistent and Universal Multi-Agent Framework for Cipher-Based Jailbreak Attacks for LLMs
von: Chen, Boyuan, et al.
Veröffentlicht: (2025)
von: Chen, Boyuan, et al.
Veröffentlicht: (2025)
Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical Evidence
von: Fu, Shaopeng, et al.
Veröffentlicht: (2025)
von: Fu, Shaopeng, et al.
Veröffentlicht: (2025)
A Causal Perspective for Enhancing Jailbreak Attack and Defense
von: Pan, Licheng, et al.
Veröffentlicht: (2026)
von: Pan, Licheng, et al.
Veröffentlicht: (2026)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
von: Chao, Patrick, et al.
Veröffentlicht: (2024)
von: Chao, Patrick, et al.
Veröffentlicht: (2024)
Noise as a Probe: Membership Inference Attacks on Diffusion Models Leveraging Initial Noise
von: Lian, Puwei, et al.
Veröffentlicht: (2026)
von: Lian, Puwei, et al.
Veröffentlicht: (2026)
Towards Automatic Hands-on-Keyboard Attack Detection Using LLMs in EDR Solutions
von: Portnoy, Amit, et al.
Veröffentlicht: (2024)
von: Portnoy, Amit, et al.
Veröffentlicht: (2024)
Representing LLMs in Prompt Semantic Task Space
von: Kashani, Idan, et al.
Veröffentlicht: (2025)
von: Kashani, Idan, et al.
Veröffentlicht: (2025)
SoK: Reducing the Vulnerability of Fine-tuned Language Models to Membership Inference Attacks
von: Amit, Guy, et al.
Veröffentlicht: (2024)
von: Amit, Guy, et al.
Veröffentlicht: (2024)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks
von: Nurlanov, Zhakshylyk, et al.
Veröffentlicht: (2026)
von: Nurlanov, Zhakshylyk, et al.
Veröffentlicht: (2026)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
von: Hu, Hanjiang, et al.
Veröffentlicht: (2025)
von: Hu, Hanjiang, et al.
Veröffentlicht: (2025)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
von: Li, Qizhang, et al.
Veröffentlicht: (2024)
von: Li, Qizhang, et al.
Veröffentlicht: (2024)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
von: Wu, Yuanwei, et al.
Veröffentlicht: (2023)
von: Wu, Yuanwei, et al.
Veröffentlicht: (2023)
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
von: Kim, Heegyu, et al.
Veröffentlicht: (2024)
von: Kim, Heegyu, et al.
Veröffentlicht: (2024)
AMED: Automatic Mixed-Precision Quantization for Edge Devices
von: Kimhi, Moshe, et al.
Veröffentlicht: (2022)
von: Kimhi, Moshe, et al.
Veröffentlicht: (2022)
Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
von: Yang, Junxiao, et al.
Veröffentlicht: (2025)
von: Yang, Junxiao, et al.
Veröffentlicht: (2025)
From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
von: Wang, Zi, et al.
Veröffentlicht: (2024)
von: Wang, Zi, et al.
Veröffentlicht: (2024)
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
von: Cornacchia, Giandomenico, et al.
Veröffentlicht: (2024)
von: Cornacchia, Giandomenico, et al.
Veröffentlicht: (2024)
Jailbreaking Large Language Models in Infinitely Many Ways
von: Goldstein, Oliver, et al.
Veröffentlicht: (2025)
von: Goldstein, Oliver, et al.
Veröffentlicht: (2025)
JULI: Jailbreak Large Language Models by Self-Introspection
von: Wang, Jesson, et al.
Veröffentlicht: (2025)
von: Wang, Jesson, et al.
Veröffentlicht: (2025)
DeepInception: Hypnotize Large Language Model to Be Jailbreaker
von: Li, Xuan, et al.
Veröffentlicht: (2023)
von: Li, Xuan, et al.
Veröffentlicht: (2023)
Defending Jailbreak Prompts via In-Context Adversarial Game
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)
Adaptive Probe-based Steering for Robust LLM Jailbreaking
von: Chen, Junxi, et al.
Veröffentlicht: (2026)
von: Chen, Junxi, et al.
Veröffentlicht: (2026)
LLMStinger: Jailbreaking LLMs using RL fine-tuned LLMs
von: Jha, Piyush, et al.
Veröffentlicht: (2024)
von: Jha, Piyush, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Sparse patches adversarial attacks via extrapolating point-wise information
von: Nemcovsky, Yaniv, et al.
Veröffentlicht: (2024) -
Silenced Biases: The Dark Side LLMs Learned to Refuse
von: Himelstein, Rom, et al.
Veröffentlicht: (2025) -
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models
von: Rahimi, Eliron, et al.
Veröffentlicht: (2026) -
Model Inversion Attacks Meet Cryptographic Fuzzy Extractors
von: Prabhakar, Mallika, et al.
Veröffentlicht: (2025) -
Indiscriminate Data Poisoning Attacks on Pre-trained Feature Extractors
von: Lu, Yiwei, et al.
Veröffentlicht: (2024)