Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kulshreshtha, Devang, Su, Hang, Jin, Haibo, Hegde, Chinmay, Wang, Haohan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Now You Hear Me: Audio Narrative Attacks Against Large Audio-Language Models
von: Yu, Ye, et al.
Veröffentlicht: (2026)
von: Yu, Ye, et al.
Veröffentlicht: (2026)
Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs
von: Jin, Haibo, et al.
Veröffentlicht: (2025)
von: Jin, Haibo, et al.
Veröffentlicht: (2025)
STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents
von: Li, Jing-Jing, et al.
Veröffentlicht: (2025)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2025)
InfoFlood: Jailbreaking Large Language Models with Information Overload
von: Yadav, Advait, et al.
Veröffentlicht: (2025)
von: Yadav, Advait, et al.
Veröffentlicht: (2025)
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models
von: Jin, Haibo, et al.
Veröffentlicht: (2025)
von: Jin, Haibo, et al.
Veröffentlicht: (2025)
The Subtle Art of Defection: Understanding Uncooperative Behaviors in LLM based Multi-Agent Systems
von: Kulshreshtha, Devang, et al.
Veröffentlicht: (2025)
von: Kulshreshtha, Devang, et al.
Veröffentlicht: (2025)
GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language Models
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
Reasoning Can Hurt the Inductive Abilities of Large Language Models
von: Jin, Haibo, et al.
Veröffentlicht: (2025)
von: Jin, Haibo, et al.
Veröffentlicht: (2025)
JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
von: Zhang, Peiyan, et al.
Veröffentlicht: (2025)
von: Zhang, Peiyan, et al.
Veröffentlicht: (2025)
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
von: Zhou, Andy, et al.
Veröffentlicht: (2024)
von: Zhou, Andy, et al.
Veröffentlicht: (2024)
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
WILDCHAT-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training
von: Feuer, Benjamin, et al.
Veröffentlicht: (2025)
von: Feuer, Benjamin, et al.
Veröffentlicht: (2025)
TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization
von: Fu, Lucheng, et al.
Veröffentlicht: (2026)
von: Fu, Lucheng, et al.
Veröffentlicht: (2026)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
Sequential Editing for Lifelong Training of Speech Recognition Models
von: Kulshreshtha, Devang, et al.
Veröffentlicht: (2024)
von: Kulshreshtha, Devang, et al.
Veröffentlicht: (2024)
Headlines You Won't Forget: Can Pronoun Insertion Increase Memorability?
von: Meyer, Selina, et al.
Veröffentlicht: (2026)
von: Meyer, Selina, et al.
Veröffentlicht: (2026)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)
Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain
von: Cheng, Gang, et al.
Veröffentlicht: (2025)
von: Cheng, Gang, et al.
Veröffentlicht: (2025)
You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
von: Cao, Bochuan, et al.
Veröffentlicht: (2025)
von: Cao, Bochuan, et al.
Veröffentlicht: (2025)
`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts
von: Schoene, Annika M, et al.
Veröffentlicht: (2025)
von: Schoene, Annika M, et al.
Veröffentlicht: (2025)
Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs
von: Hartmann, David, et al.
Veröffentlicht: (2026)
von: Hartmann, David, et al.
Veröffentlicht: (2026)
Unveiling the Lexical Sensitivity of LLMs: Combinatorial Optimization for Prompt Enhancement
von: Zhan, Pengwei, et al.
Veröffentlicht: (2024)
von: Zhan, Pengwei, et al.
Veröffentlicht: (2024)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
von: Chen, Yunhao, et al.
Veröffentlicht: (2025)
von: Chen, Yunhao, et al.
Veröffentlicht: (2025)
Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
von: Yu, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Yu, Zhiyuan, et al.
Veröffentlicht: (2024)
PARDEN, Can You Repeat That? Defending against Jailbreaks via Repetition
von: Zhang, Ziyang, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2024)
Catch Me If You Can? Not Yet: LLMs Still Struggle to Imitate the Implicit Writing Styles of Everyday Authors
von: Wang, Zhengxiang, et al.
Veröffentlicht: (2025)
von: Wang, Zhengxiang, et al.
Veröffentlicht: (2025)
Do Self-Evolving Agents Forget? Capability Degradation and Preservation in Lifelong LLM Agent Adaptation
von: Yu, Ye, et al.
Veröffentlicht: (2026)
von: Yu, Ye, et al.
Veröffentlicht: (2026)
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2023)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2023)
Trust Me, I Can Convince You: The Contextualized Argument Appraisal Framework
von: Greschner, Lynn, et al.
Veröffentlicht: (2025)
von: Greschner, Lynn, et al.
Veröffentlicht: (2025)
Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
von: Sedláček, Šimon, et al.
Veröffentlicht: (2025)
von: Sedláček, Šimon, et al.
Veröffentlicht: (2025)
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
von: Xie, Yueqi, et al.
Veröffentlicht: (2024)
von: Xie, Yueqi, et al.
Veröffentlicht: (2024)
Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?
von: Bhaskar, Adithya, et al.
Veröffentlicht: (2025)
von: Bhaskar, Adithya, et al.
Veröffentlicht: (2025)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
von: Zhang, Chiyu, et al.
Veröffentlicht: (2025)
von: Zhang, Chiyu, et al.
Veröffentlicht: (2025)
Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs
von: Li, Xiaoxia, et al.
Veröffentlicht: (2024)
von: Li, Xiaoxia, et al.
Veröffentlicht: (2024)
Jailbreaking with Universal Multi-Prompts
von: Hsu, Yu-Ling, et al.
Veröffentlicht: (2025)
von: Hsu, Yu-Ling, et al.
Veröffentlicht: (2025)
Catch Me If You Can Describe Me: Open-Vocabulary Camouflaged Instance Segmentation with Diffusion
von: Vu, Tuan-Anh, et al.
Veröffentlicht: (2023)
von: Vu, Tuan-Anh, et al.
Veröffentlicht: (2023)
Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
von: Joo, Seongho, et al.
Veröffentlicht: (2025)
von: Joo, Seongho, et al.
Veröffentlicht: (2025)
Iterative Self-Tuning LLMs for Enhanced Jailbreaking Capabilities
von: Sun, Chung-En, et al.
Veröffentlicht: (2024)
von: Sun, Chung-En, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Now You Hear Me: Audio Narrative Attacks Against Large Audio-Language Models
von: Yu, Ye, et al.
Veröffentlicht: (2026) -
Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters
von: Jin, Haibo, et al.
Veröffentlicht: (2024) -
GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs
von: Jin, Haibo, et al.
Veröffentlicht: (2025) -
STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents
von: Li, Jing-Jing, et al.
Veröffentlicht: (2025) -
InfoFlood: Jailbreaking Large Language Models with Information Overload
von: Yadav, Advait, et al.
Veröffentlicht: (2025)