Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nurlanov, Zhakshylyk, Schmidt, Frank R., Bernard, Florian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
von: Yang, Junxiao, et al.
Veröffentlicht: (2025)
von: Yang, Junxiao, et al.
Veröffentlicht: (2025)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
von: Nikolić, Kristina, et al.
Veröffentlicht: (2025)
von: Nikolić, Kristina, et al.
Veröffentlicht: (2025)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
A Causal Perspective for Enhancing Jailbreak Attack and Defense
von: Pan, Licheng, et al.
Veröffentlicht: (2026)
von: Pan, Licheng, et al.
Veröffentlicht: (2026)
Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning
von: Hasan, Adib, et al.
Veröffentlicht: (2024)
von: Hasan, Adib, et al.
Veröffentlicht: (2024)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025)
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025)
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
von: Wu, Yuanwei, et al.
Veröffentlicht: (2023)
von: Wu, Yuanwei, et al.
Veröffentlicht: (2023)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs
von: Chen, Wenyu, et al.
Veröffentlicht: (2026)
von: Chen, Wenyu, et al.
Veröffentlicht: (2026)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
von: Lin, Shi, et al.
Veröffentlicht: (2024)
von: Lin, Shi, et al.
Veröffentlicht: (2024)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
von: Wang, Zi, et al.
Veröffentlicht: (2024)
von: Wang, Zi, et al.
Veröffentlicht: (2024)
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
von: Cornacchia, Giandomenico, et al.
Veröffentlicht: (2024)
von: Cornacchia, Giandomenico, et al.
Veröffentlicht: (2024)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
von: Rando, Javier, et al.
Veröffentlicht: (2024)
von: Rando, Javier, et al.
Veröffentlicht: (2024)
Jailbreaking LLMs via Calibration
von: Lu, Yuxuan, et al.
Veröffentlicht: (2026)
von: Lu, Yuxuan, et al.
Veröffentlicht: (2026)
Optimal Defenses Against Gradient Reconstruction Attacks
von: Chen, Yuxiao, et al.
Veröffentlicht: (2024)
von: Chen, Yuxiao, et al.
Veröffentlicht: (2024)
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2024)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2024)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2024)
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2024)
Practical Feasibility of Gradient Inversion Attacks in Federated Learning
von: Valadi, Viktor, et al.
Veröffentlicht: (2025)
von: Valadi, Viktor, et al.
Veröffentlicht: (2025)
Geminio: Language-Guided Gradient Inversion Attacks in Federated Learning
von: Shan, Junjie, et al.
Veröffentlicht: (2024)
von: Shan, Junjie, et al.
Veröffentlicht: (2024)
Universal Jailbreak Backdoors from Poisoned Human Feedback
von: Rando, Javier, et al.
Veröffentlicht: (2023)
von: Rando, Javier, et al.
Veröffentlicht: (2023)
Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
von: Guo, Qiming, et al.
Veröffentlicht: (2025)
von: Guo, Qiming, et al.
Veröffentlicht: (2025)
Mitigating Many-Shot Jailbreaking
von: Ackerman, Christopher M., et al.
Veröffentlicht: (2025)
von: Ackerman, Christopher M., et al.
Veröffentlicht: (2025)
No More Guessing: a Verifiable Gradient Inversion Attack in Federated Learning
von: Diana, Francesco, et al.
Veröffentlicht: (2026)
von: Diana, Francesco, et al.
Veröffentlicht: (2026)
A Numerical Gradient Inversion Attack in Variational Quantum Neural-Networks
von: Papadopoulos, Georgios, et al.
Veröffentlicht: (2025)
von: Papadopoulos, Georgios, et al.
Veröffentlicht: (2025)
Beyond Gradient and Priors in Privacy Attacks: Leveraging Pooler Layer Inputs of Language Models in Federated Learning
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
Attention Tracker: Detecting Prompt Injection Attacks in LLMs
von: Hung, Kuo-Han, et al.
Veröffentlicht: (2024)
von: Hung, Kuo-Han, et al.
Veröffentlicht: (2024)
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
von: Fang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Fang, Zhicheng, et al.
Veröffentlicht: (2026)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean Reference
von: Li, Jianwei, et al.
Veröffentlicht: (2026)
von: Li, Jianwei, et al.
Veröffentlicht: (2026)
DR-Encoder: Encode Low-rank Gradients with Random Prior for Large Language Models Differentially Privately
von: Wu, Huiwen, et al.
Veröffentlicht: (2024)
von: Wu, Huiwen, et al.
Veröffentlicht: (2024)
Learning to Defend by Attacking (and Vice-Versa): Transfer of Learning in Cybersecurity Games
von: Malloy, Tailia, et al.
Veröffentlicht: (2023)
von: Malloy, Tailia, et al.
Veröffentlicht: (2023)
GI-PIP: Do We Require Impractical Auxiliary Dataset for Gradient Inversion Attacks?
von: Sun, Yu, et al.
Veröffentlicht: (2024)
von: Sun, Yu, et al.
Veröffentlicht: (2024)
AGSOA:Graph Neural Network Targeted Attack Based on Average Gradient and Structure Optimization
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
von: Peng, Benji, et al.
Veröffentlicht: (2024)
von: Peng, Benji, et al.
Veröffentlicht: (2024)
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
von: Panfilov, Alexander, et al.
Veröffentlicht: (2026)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
von: Yang, Junxiao, et al.
Veröffentlicht: (2025) -
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024) -
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024) -
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
von: Nikolić, Kristina, et al.
Veröffentlicht: (2025) -
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)