LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Shi, Yang, Hongming, Li, Rongchang, Wang, Xun, Lin, Changting, Xing, Wenpeng, Han, Meng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks
von: Li, Rongchang, et al.
Veröffentlicht: (2024)
von: Li, Rongchang, et al.
Veröffentlicht: (2024)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
von: Yu, Miao, et al.
Veröffentlicht: (2024)
von: Yu, Miao, et al.
Veröffentlicht: (2024)
Distract Large Language Models for Automatic Jailbreak Attack
von: Xiao, Zeguan, et al.
Veröffentlicht: (2024)
von: Xiao, Zeguan, et al.
Veröffentlicht: (2024)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
StructuralSleight: Automated Jailbreak Attacks on Large Language Models Utilizing Uncommon Text-Organization Structures
von: Li, Bangxin, et al.
Veröffentlicht: (2024)
von: Li, Bangxin, et al.
Veröffentlicht: (2024)
Harnessing Task Overload for Scalable Jailbreak Attacks on Large Language Models
von: Dong, Yiting, et al.
Veröffentlicht: (2024)
von: Dong, Yiting, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models
von: Yang, Guangyu, et al.
Veröffentlicht: (2025)
von: Yang, Guangyu, et al.
Veröffentlicht: (2025)
EverTracer: Hunting Stolen Large Language Models via Stealthy and Robust Probabilistic Fingerprint
von: Xu, Zhenhua, et al.
Veröffentlicht: (2025)
von: Xu, Zhenhua, et al.
Veröffentlicht: (2025)
Jailbreaking Leaves a Trace: Understanding and Detecting Jailbreak Attacks from Internal Representations of Large Language Models
von: Kadali, Sri Durga Sai Sowmya, et al.
Veröffentlicht: (2026)
von: Kadali, Sri Durga Sai Sowmya, et al.
Veröffentlicht: (2026)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models
von: Xu, Zihao, et al.
Veröffentlicht: (2024)
von: Xu, Zihao, et al.
Veröffentlicht: (2024)
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
von: Feng, Yingchaojie, et al.
Veröffentlicht: (2024)
von: Feng, Yingchaojie, et al.
Veröffentlicht: (2024)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
von: Chen, Yunhao, et al.
Veröffentlicht: (2025)
von: Chen, Yunhao, et al.
Veröffentlicht: (2025)
Denial-of-Service Poisoning Attacks against Large Language Models
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
von: You, Wenhao, et al.
Veröffentlicht: (2025)
von: You, Wenhao, et al.
Veröffentlicht: (2025)
CCJA: Context-Coherent Jailbreak Attack for Aligned Large Language Models
von: Zhou, Guanghao, et al.
Veröffentlicht: (2025)
von: Zhou, Guanghao, et al.
Veröffentlicht: (2025)
Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
von: Jia, Xiaojun, et al.
Veröffentlicht: (2024)
von: Jia, Xiaojun, et al.
Veröffentlicht: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
von: Xu, Zhao, et al.
Veröffentlicht: (2024)
von: Xu, Zhao, et al.
Veröffentlicht: (2024)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
Large Reasoning Models Are Autonomous Jailbreak Agents
von: Hagendorff, Thilo, et al.
Veröffentlicht: (2025)
von: Hagendorff, Thilo, et al.
Veröffentlicht: (2025)
InfoFlood: Jailbreaking Large Language Models with Information Overload
von: Yadav, Advait, et al.
Veröffentlicht: (2025)
von: Yadav, Advait, et al.
Veröffentlicht: (2025)
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
von: Liu, Yanjiang, et al.
Veröffentlicht: (2025)
von: Liu, Yanjiang, et al.
Veröffentlicht: (2025)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
von: Liu, Fan, et al.
Veröffentlicht: (2024)
von: Liu, Fan, et al.
Veröffentlicht: (2024)
STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents
von: Li, Jing-Jing, et al.
Veröffentlicht: (2025)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2025)
QueryAttack: Jailbreaking Aligned Large Language Models Using Structured Non-natural Query Language
von: Zou, Qingsong, et al.
Veröffentlicht: (2025)
von: Zou, Qingsong, et al.
Veröffentlicht: (2025)
E-SAGE: Explainability-based Defense Against Backdoor Attacks on Graph Neural Networks
von: Yuan, Dingqiang, et al.
Veröffentlicht: (2024)
von: Yuan, Dingqiang, et al.
Veröffentlicht: (2024)
Prompt Stealing Attacks Against Large Language Models
von: Sha, Zeyang, et al.
Veröffentlicht: (2024)
von: Sha, Zeyang, et al.
Veröffentlicht: (2024)
Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks
von: Zhao, Jiawei, et al.
Veröffentlicht: (2024)
von: Zhao, Jiawei, et al.
Veröffentlicht: (2024)
RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2024)
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2024)
MCP-Guard: A Multi-Stage Defense-in-Depth Framework for Securing Model Context Protocol in Agentic AI
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
von: Akbar-Tajari, Mohammad, et al.
Veröffentlicht: (2025)
von: Akbar-Tajari, Mohammad, et al.
Veröffentlicht: (2025)
Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
von: Pu, Rui, et al.
Veröffentlicht: (2024)
von: Pu, Rui, et al.
Veröffentlicht: (2024)
AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
von: Lv, Lijia, et al.
Veröffentlicht: (2024)
von: Lv, Lijia, et al.
Veröffentlicht: (2024)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
von: Yin, Ziyi, et al.
Veröffentlicht: (2025)
von: Yin, Ziyi, et al.
Veröffentlicht: (2025)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025) -
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
von: Wang, Yidan, et al.
Veröffentlicht: (2025) -
GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks
von: Li, Rongchang, et al.
Veröffentlicht: (2024) -
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
von: Yu, Miao, et al.
Veröffentlicht: (2024) -
Distract Large Language Models for Automatic Jailbreak Attack
von: Xiao, Zeguan, et al.
Veröffentlicht: (2024)