Plentiful Jailbreaks with String Compositions
Fuente:
arXiv
Saved in:
| Main Author: | Huang, Brian R. Y. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Endless Jailbreaks with Bijection Learning
by: Huang, Brian R. Y., et al.
Published: (2024)
by: Huang, Brian R. Y., et al.
Published: (2024)
Jailbreaking to Jailbreak
by: Kritz, Jeremy, et al.
Published: (2025)
by: Kritz, Jeremy, et al.
Published: (2025)
Don't Say No: Jailbreaking LLM by Suppressing Refusal
by: Zhou, Yukai, et al.
Published: (2024)
by: Zhou, Yukai, et al.
Published: (2024)
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
by: Zhou, Weikang, et al.
Published: (2024)
by: Zhou, Weikang, et al.
Published: (2024)
xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking
by: Lee, Sunbowen, et al.
Published: (2025)
by: Lee, Sunbowen, et al.
Published: (2025)
Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
by: Ji, Haoxuan, et al.
Published: (2024)
by: Ji, Haoxuan, et al.
Published: (2024)
StringLLM: Understanding the String Processing Capability of Large Language Models
by: Wang, Xilong, et al.
Published: (2024)
by: Wang, Xilong, et al.
Published: (2024)
Jailbreaking? One Step Is Enough!
by: Zheng, Weixiong, et al.
Published: (2024)
by: Zheng, Weixiong, et al.
Published: (2024)
Rewrite to Jailbreak: Discover Learnable and Transferable Implicit Harmfulness Instruction
by: Huang, Yuting, et al.
Published: (2025)
by: Huang, Yuting, et al.
Published: (2025)
DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Can Large Language Models Automatically Jailbreak GPT-4V?
by: Wu, Yuanwei, et al.
Published: (2024)
by: Wu, Yuanwei, et al.
Published: (2024)
Enhancing Jailbreak Attacks with Diversity Guidance
by: Zhang, Xu, et al.
Published: (2024)
by: Zhang, Xu, et al.
Published: (2024)
Persona Jailbreaking in Large Language Models
by: Sandhan, Jivnesh, et al.
Published: (2026)
by: Sandhan, Jivnesh, et al.
Published: (2026)
ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025)
by: Murphy, Brendan, et al.
Published: (2025)
Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
by: Zhang, Zhexin, et al.
Published: (2023)
by: Zhang, Zhexin, et al.
Published: (2023)
LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds
by: Beetham, James, et al.
Published: (2024)
by: Beetham, James, et al.
Published: (2024)
Many-Turn Jailbreaking
by: Yang, Xianjun, et al.
Published: (2025)
by: Yang, Xianjun, et al.
Published: (2025)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
by: Liu, Songyang, et al.
Published: (2026)
by: Liu, Songyang, et al.
Published: (2026)
Jailbreaking Large Language Models Through Alignment Vulnerabilities in Out-of-Distribution Settings
by: Huang, Yue, et al.
Published: (2024)
by: Huang, Yue, et al.
Published: (2024)
Weak-to-Strong Jailbreaking on Large Language Models
by: Zhao, Xuandong, et al.
Published: (2024)
by: Zhao, Xuandong, et al.
Published: (2024)
[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs
by: Rao, Abhinav, et al.
Published: (2024)
by: Rao, Abhinav, et al.
Published: (2024)
Diversity Helps Jailbreak Large Language Models
by: Zhao, Weiliang, et al.
Published: (2024)
by: Zhao, Weiliang, et al.
Published: (2024)
Multilingual Jailbreak Challenges in Large Language Models
by: Deng, Yue, et al.
Published: (2023)
by: Deng, Yue, et al.
Published: (2023)
Jailbreaking Large Language Models with Morality Attacks
by: Su, Ying, et al.
Published: (2026)
by: Su, Ying, et al.
Published: (2026)
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
by: Feng, Yingchaojie, et al.
Published: (2024)
by: Feng, Yingchaojie, et al.
Published: (2024)
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
by: Rao, Abhinav, et al.
Published: (2023)
by: Rao, Abhinav, et al.
Published: (2023)
RAID: Refusal-Aware and Integrated Decoding for Jailbreaking LLMs
by: Nguyen, Tuan T., et al.
Published: (2025)
by: Nguyen, Tuan T., et al.
Published: (2025)
ASCIIEval: Benchmarking Models' Visual Perception in Text Strings via ASCII Art
by: Jia, Qi, et al.
Published: (2024)
by: Jia, Qi, et al.
Published: (2024)
Attention-Aware GNN-based Input Defense against Multi-Turn LLM Jailbreak
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
GRAF: Multi-turn Jailbreaking via Global Refinement and Active Fabrication
by: Tang, Hua, et al.
Published: (2025)
by: Tang, Hua, et al.
Published: (2025)
Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs
by: Li, Xiaoxia, et al.
Published: (2024)
by: Li, Xiaoxia, et al.
Published: (2024)
AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation
by: Wang, Zijun, et al.
Published: (2024)
by: Wang, Zijun, et al.
Published: (2024)
Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
by: Peng, Alwin, et al.
Published: (2024)
by: Peng, Alwin, et al.
Published: (2024)
A Troublemaker with Contagious Jailbreak Makes Chaos in Honest Towns
by: Men, Tianyi, et al.
Published: (2024)
by: Men, Tianyi, et al.
Published: (2024)
Intention Analysis Makes LLMs A Good Jailbreak Defender
by: Zhang, Yuqi, et al.
Published: (2024)
by: Zhang, Yuqi, et al.
Published: (2024)
Causal Front-Door Adjustment for Robust Jailbreak Attacks on LLMs
by: Zhou, Yao, et al.
Published: (2026)
by: Zhou, Yao, et al.
Published: (2026)
Do Internal Layers of LLMs Reveal Patterns for Jailbreak Detection?
by: Kadali, Sri Durga Sai Sowmya, et al.
Published: (2025)
by: Kadali, Sri Durga Sai Sowmya, et al.
Published: (2025)
Dialogue Injection Attack: Jailbreaking LLMs through Context Manipulation
by: Meng, Wenlong, et al.
Published: (2025)
by: Meng, Wenlong, et al.
Published: (2025)
Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
by: Atil, Berk, et al.
Published: (2025)
by: Atil, Berk, et al.
Published: (2025)
Similar Items
-
Endless Jailbreaks with Bijection Learning
by: Huang, Brian R. Y., et al.
Published: (2024) -
Jailbreaking to Jailbreak
by: Kritz, Jeremy, et al.
Published: (2025) -
Don't Say No: Jailbreaking LLM by Suppressing Refusal
by: Zhou, Yukai, et al.
Published: (2024) -
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
by: Zhou, Weikang, et al.
Published: (2024) -
xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking
by: Lee, Sunbowen, et al.
Published: (2025)