RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Xuan, Nie, Yuzhou, Yan, Lu, Mao, Yunshu, Guo, Wenbo, Zhang, Xiangyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided Search
by: Chen, Xuan, et al.
Published: (2024)
by: Chen, Xuan, et al.
Published: (2024)
Temporal Logic-Based Multi-Vehicle Backdoor Attacks against Offline RL Agents in End-to-end Autonomous Driving
by: Chen, Xuan, et al.
Published: (2025)
by: Chen, Xuan, et al.
Published: (2025)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
by: Xiong, Chen, et al.
Published: (2024)
by: Xiong, Chen, et al.
Published: (2024)
KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
by: Liu, Shuyuan, et al.
Published: (2025)
by: Liu, Shuyuan, et al.
Published: (2025)
Black-box Membership Inference Attacks against Fine-tuned Diffusion Models
by: Pang, Yan, et al.
Published: (2023)
by: Pang, Yan, et al.
Published: (2023)
Query Provenance Analysis: Efficient and Robust Defense against Query-based Black-box Attacks
by: Li, Shaofei, et al.
Published: (2024)
by: Li, Shaofei, et al.
Published: (2024)
JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs
by: Li, Hongyi, et al.
Published: (2024)
by: Li, Hongyi, et al.
Published: (2024)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
by: Wang, Xinkai, et al.
Published: (2025)
by: Wang, Xinkai, et al.
Published: (2025)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
by: Chen, Bocheng, et al.
Published: (2024)
by: Chen, Bocheng, et al.
Published: (2024)
White-box Membership Inference Attacks against Diffusion Models
by: Pang, Yan, et al.
Published: (2023)
by: Pang, Yan, et al.
Published: (2023)
Multi-task Adversarial Attacks against Black-box Model with Few-shot Queries
by: Wang, Wenqiang, et al.
Published: (2025)
by: Wang, Wenqiang, et al.
Published: (2025)
Online Poisoning Attack Against Reinforcement Learning under Black-box Environments
by: Li, Jianhui, et al.
Published: (2024)
by: Li, Jianhui, et al.
Published: (2024)
Backdoor Attacks against Hybrid Classical-Quantum Neural Networks
by: Guo, Ji, et al.
Published: (2024)
by: Guo, Ji, et al.
Published: (2024)
BlackboxBench: A Comprehensive Benchmark of Black-box Adversarial Attacks
by: Zheng, Meixi, et al.
Published: (2023)
by: Zheng, Meixi, et al.
Published: (2023)
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
by: Wang, Youze, et al.
Published: (2025)
by: Wang, Youze, et al.
Published: (2025)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
by: Lu, Lin, et al.
Published: (2024)
by: Lu, Lin, et al.
Published: (2024)
Activation Surgery: Jailbreaking White-box LLMs without Touching the Prompt
by: Jenny, Maël, et al.
Published: (2026)
by: Jenny, Maël, et al.
Published: (2026)
TrojFM: Resource-efficient Backdoor Attacks against Very Large Foundation Models
by: Nie, Yuzhou., et al.
Published: (2024)
by: Nie, Yuzhou., et al.
Published: (2024)
Multi-granular Adversarial Attacks against Black-box Neural Ranking Models
by: Liu, Yu-An, et al.
Published: (2024)
by: Liu, Yu-An, et al.
Published: (2024)
Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models
by: Chen, Zeyuan, et al.
Published: (2026)
by: Chen, Zeyuan, et al.
Published: (2026)
AISA: Awakening Intrinsic Safety Awareness in Large Language Models against Jailbreak Attacks
by: Song, Weiming, et al.
Published: (2026)
by: Song, Weiming, et al.
Published: (2026)
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
by: Cui, Tiehan, et al.
Published: (2025)
by: Cui, Tiehan, et al.
Published: (2025)
Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
by: Zhong, Xingwei, et al.
Published: (2025)
by: Zhong, Xingwei, et al.
Published: (2025)
LLMStinger: Jailbreaking LLMs using RL fine-tuned LLMs
by: Jha, Piyush, et al.
Published: (2024)
by: Jha, Piyush, et al.
Published: (2024)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
by: Nie, Yuzhou, et al.
Published: (2024)
by: Nie, Yuzhou, et al.
Published: (2024)
ASPIRER: Bypassing System Prompts With Permutation-based Backdoors in LLMs
by: Yan, Lu, et al.
Published: (2024)
by: Yan, Lu, et al.
Published: (2024)
From Head to Tail: Efficient Black-box Model Inversion Attack via Long-tailed Learning
by: Li, Ziang, et al.
Published: (2025)
by: Li, Ziang, et al.
Published: (2025)
BETA: Automated Black-box Exploration for Timing Attacks in Processors
by: Chen, Congcong, et al.
Published: (2024)
by: Chen, Congcong, et al.
Published: (2024)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
by: Chen, Yunhao, et al.
Published: (2025)
by: Chen, Yunhao, et al.
Published: (2025)
From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem
by: Mao, Yanxu, et al.
Published: (2025)
by: Mao, Yanxu, et al.
Published: (2025)
Beyond Surface-Level Patterns: An Essence-Driven Defense Framework Against Jailbreak Attacks in LLMs
by: Xiang, Shiyu, et al.
Published: (2025)
by: Xiang, Shiyu, et al.
Published: (2025)
FlipAttack: Jailbreak LLMs via Flipping
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
by: Zhang, Chiyu, et al.
Published: (2025)
by: Zhang, Chiyu, et al.
Published: (2025)
MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
by: Chen, Tailun, et al.
Published: (2025)
by: Chen, Tailun, et al.
Published: (2025)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
by: Akbar-Tajari, Mohammad, et al.
Published: (2025)
by: Akbar-Tajari, Mohammad, et al.
Published: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
by: Ma, Jiachen, et al.
Published: (2024)
by: Ma, Jiachen, et al.
Published: (2024)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
by: Chen, Yulin, et al.
Published: (2024)
by: Chen, Yulin, et al.
Published: (2024)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance
by: Zhang, Zaixi, et al.
Published: (2025)
by: Zhang, Zaixi, et al.
Published: (2025)
Similar Items
-
When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided Search
by: Chen, Xuan, et al.
Published: (2024) -
Temporal Logic-Based Multi-Vehicle Backdoor Attacks against Offline RL Agents in End-to-end Autonomous Driving
by: Chen, Xuan, et al.
Published: (2025) -
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
by: Xiong, Chen, et al.
Published: (2024) -
KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
by: Liu, Shuyuan, et al.
Published: (2025) -
Black-box Membership Inference Attacks against Fine-tuned Diffusion Models
by: Pang, Yan, et al.
Published: (2023)