Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhaoqi, He, Daqing, Zhang, Zijian, Li, Xin, Zhu, Liehuang, Li, Meng, Liu, Jiamou |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jailbreaking Large Language Models through Iterative Tool-Disguised Attacks via Reinforcement Learning
by: Wang, Zhaoqi, et al.
Published: (2026)
by: Wang, Zhaoqi, et al.
Published: (2026)
PathSeeker: Exploring LLM Security Vulnerabilities with a Reinforcement Learning-Based Jailbreak Approach
by: Lin, Zhihao, et al.
Published: (2024)
by: Lin, Zhihao, et al.
Published: (2024)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
Involuntary Jailbreak: On Self-Prompting Attacks
by: Guo, Yangyang, et al.
Published: (2025)
by: Guo, Yangyang, et al.
Published: (2025)
OptiLeak: Efficient Prompt Reconstruction via Reinforcement Learning in Multi-tenant LLM Services
by: Wang, Longxiang, et al.
Published: (2026)
by: Wang, Longxiang, et al.
Published: (2026)
Large Language Model-driven Security Assistant for Internet of Things via Chain-of-Thought
by: Zeng, Mingfei, et al.
Published: (2025)
by: Zeng, Mingfei, et al.
Published: (2025)
TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
by: Yoon, Sung-Hoon, et al.
Published: (2026)
by: Yoon, Sung-Hoon, et al.
Published: (2026)
Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks
by: Zhang, Yingjie, et al.
Published: (2025)
by: Zhang, Yingjie, et al.
Published: (2025)
Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks
by: Geng, Jianing, et al.
Published: (2025)
by: Geng, Jianing, et al.
Published: (2025)
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
by: Li, Xirui, et al.
Published: (2024)
by: Li, Xirui, et al.
Published: (2024)
Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs
by: Ling, Zijian, et al.
Published: (2026)
by: Ling, Zijian, et al.
Published: (2026)
MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
by: Zhu, Wentian, et al.
Published: (2025)
by: Zhu, Wentian, et al.
Published: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
by: Ma, Jiachen, et al.
Published: (2024)
by: Ma, Jiachen, et al.
Published: (2024)
FlipAttack: Jailbreak LLMs via Flipping
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search
by: Huang, Xun, et al.
Published: (2026)
by: Huang, Xun, et al.
Published: (2026)
OMNISEC: LLM-Driven Provenance-based Intrusion Detection via Retrieval-Augmented Behavior Prompting
by: Cheng, Wenrui, et al.
Published: (2025)
by: Cheng, Wenrui, et al.
Published: (2025)
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
by: Wu, Yuanwei, et al.
Published: (2023)
by: Wu, Yuanwei, et al.
Published: (2023)
JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path Rectification
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process
by: Wang, Peiran, et al.
Published: (2024)
by: Wang, Peiran, et al.
Published: (2024)
A Framework for Formalizing LLM Agent Security
by: Siu, Vincent, et al.
Published: (2026)
by: Siu, Vincent, et al.
Published: (2026)
A2-DIDM: Privacy-preserving Accumulator-enabled Auditing for Distributed Identity of DNN Model
by: Xie, Tianxiu, et al.
Published: (2024)
by: Xie, Tianxiu, et al.
Published: (2024)
DiffuseTrace: A Transparent and Flexible Watermarking Scheme for Latent Diffusion Model
by: Lei, Liangqi, et al.
Published: (2024)
by: Lei, Liangqi, et al.
Published: (2024)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
by: Jaiswal, Piyush, et al.
Published: (2026)
by: Jaiswal, Piyush, et al.
Published: (2026)
Preventing Non-intrusive Load Monitoring Privacy Invasion: A Precise Adversarial Attack Scheme for Networked Smart Meters
by: He, Jialing, et al.
Published: (2024)
by: He, Jialing, et al.
Published: (2024)
TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense
by: Liu, Cheng, et al.
Published: (2026)
by: Liu, Cheng, et al.
Published: (2026)
AiRacleX: Automated Detection of Price Oracle Manipulations via LLM-Driven Knowledge Mining and Prompt Generation
by: Gao, Bo, et al.
Published: (2025)
by: Gao, Bo, et al.
Published: (2025)
BESA: Boosting Encoder Stealing Attack with Perturbation Recovery
by: Ren, Xuhao, et al.
Published: (2025)
by: Ren, Xuhao, et al.
Published: (2025)
Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction
by: Liu, Tong, et al.
Published: (2024)
by: Liu, Tong, et al.
Published: (2024)
JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model
by: Nian, Yi, et al.
Published: (2025)
by: Nian, Yi, et al.
Published: (2025)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
by: Wang, Zhilong, et al.
Published: (2025)
by: Wang, Zhilong, et al.
Published: (2025)
Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation
by: Zhang, Wenhui, et al.
Published: (2025)
by: Zhang, Wenhui, et al.
Published: (2025)
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
by: Lin, Lixing, et al.
Published: (2026)
by: Lin, Lixing, et al.
Published: (2026)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
by: Guo, Weiyang, et al.
Published: (2026)
by: Guo, Weiyang, et al.
Published: (2026)
"To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
by: Sun, Zhen, et al.
Published: (2025)
by: Sun, Zhen, et al.
Published: (2025)
FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
by: Gong, Yichen, et al.
Published: (2023)
by: Gong, Yichen, et al.
Published: (2023)
The Echo Chamber Multi-Turn LLM Jailbreak
by: Alobaid, Ahmad, et al.
Published: (2026)
by: Alobaid, Ahmad, et al.
Published: (2026)
Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate
by: Qi, Senmao, et al.
Published: (2025)
by: Qi, Senmao, et al.
Published: (2025)
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
by: Saha, Shoumik, et al.
Published: (2025)
by: Saha, Shoumik, et al.
Published: (2025)
Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual Steganography
by: Li, Songze, et al.
Published: (2025)
by: Li, Songze, et al.
Published: (2025)
Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
Similar Items
-
Jailbreaking Large Language Models through Iterative Tool-Disguised Attacks via Reinforcement Learning
by: Wang, Zhaoqi, et al.
Published: (2026) -
PathSeeker: Exploring LLM Security Vulnerabilities with a Reinforcement Learning-Based Jailbreak Approach
by: Lin, Zhihao, et al.
Published: (2024) -
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
by: Zhang, Zheng, et al.
Published: (2025) -
Involuntary Jailbreak: On Self-Prompting Attacks
by: Guo, Yangyang, et al.
Published: (2025) -
OptiLeak: Efficient Prompt Reconstruction via Reinforcement Learning in Multi-tenant LLM Services
by: Wang, Longxiang, et al.
Published: (2026)