AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xiaogeng, Li, Peiran, Suh, Edward, Vorobeychik, Yevgeniy, Mao, Zhuoqing, Jha, Somesh, McDaniel, Patrick, Sun, Huan, Li, Bo, Xiao, Chaowei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
by: Liu, Xiaogeng, et al.
Published: (2025)
by: Liu, Xiaogeng, et al.
Published: (2025)
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
by: Liu, Xiaogeng, et al.
Published: (2023)
by: Liu, Xiaogeng, et al.
Published: (2023)
RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process
by: Wang, Peiran, et al.
Published: (2024)
by: Wang, Peiran, et al.
Published: (2024)
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
by: Wu, Fangzhou, et al.
Published: (2024)
by: Wu, Fangzhou, et al.
Published: (2024)
Verified Safe Reinforcement Learning for Neural Network Dynamic Models
by: Wu, Junlin, et al.
Published: (2024)
by: Wu, Junlin, et al.
Published: (2024)
ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning
by: Zhao, Zhengyue, et al.
Published: (2025)
by: Zhao, Zhengyue, et al.
Published: (2025)
CoFineLLM: Conformal Finetuning of LLMs for Language-Instructed Robot Planning
by: Wang, Jun, et al.
Published: (2025)
by: Wang, Jun, et al.
Published: (2025)
Mind the (DH) Gap! A Contrast in Risky Choices Between Reasoning and Conversational LLMs
by: Ge, Luise, et al.
Published: (2026)
by: Ge, Luise, et al.
Published: (2026)
AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety Detection
by: Luo, Weidi, et al.
Published: (2025)
by: Luo, Weidi, et al.
Published: (2025)
CyGym: A Simulation-Based Game-Theoretic Analysis Framework for Cybersecurity
by: Lanier, Michael, et al.
Published: (2025)
by: Lanier, Michael, et al.
Published: (2025)
Sliced Rényi Pufferfish Privacy: Directional Additive Noise Mechanism and Private Learning with Gradient Clipping
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
Residual-PAC Privacy: Automatic Privacy Control Beyond the Gaussian Barrier
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
A Scalable Approach to Solving Simulation-Based Network Security Games
by: Lanier, Michael, et al.
Published: (2026)
by: Lanier, Michael, et al.
Published: (2026)
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
by: Wang, Jiongxiao, et al.
Published: (2023)
by: Wang, Jiongxiao, et al.
Published: (2023)
AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
by: Luo, Weidi, et al.
Published: (2024)
by: Luo, Weidi, et al.
Published: (2024)
Preference Poisoning Attacks on Reward Model Learning
by: Wu, Junlin, et al.
Published: (2024)
by: Wu, Junlin, et al.
Published: (2024)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
by: Wang, Jiongxiao, et al.
Published: (2024)
by: Wang, Jiongxiao, et al.
Published: (2024)
Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
by: Yu, Zhiyuan, et al.
Published: (2024)
by: Yu, Zhiyuan, et al.
Published: (2024)
Learning Linear Utility Functions From Pairwise Comparison Queries
by: Ge, Luise, et al.
Published: (2024)
by: Ge, Luise, et al.
Published: (2024)
Optimized Distortion in Linear Social Choice
by: Ge, Luise, et al.
Published: (2025)
by: Ge, Luise, et al.
Published: (2025)
Online Feedback Efficient Active Target Discovery in Partially Observable Environments
by: Sarkar, Anindya, et al.
Published: (2025)
by: Sarkar, Anindya, et al.
Published: (2025)
Learning Recommender Mechanisms for Bayesian Stochastic Games
by: Guresti, Bengisu, et al.
Published: (2025)
by: Guresti, Bengisu, et al.
Published: (2025)
Adversarial Reinforcement Learning for Detecting False Data Injection Attacks in Vehicular Routing
by: Eghtesad, Taha, et al.
Published: (2026)
by: Eghtesad, Taha, et al.
Published: (2026)
Active Target Discovery under Uninformative Prior: The Power of Permanent and Transient Memory
by: Sarkar, Anindya, et al.
Published: (2025)
by: Sarkar, Anindya, et al.
Published: (2025)
Multi-Agent Reinforcement Learning for Assessing False-Data Injection Attacks on Transportation Networks
by: Eghtesad, Taha, et al.
Published: (2023)
by: Eghtesad, Taha, et al.
Published: (2023)
ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention
by: Wang, Xinyan, et al.
Published: (2026)
by: Wang, Xinyan, et al.
Published: (2026)
MetaAgent: Automatically Constructing Multi-Agent Systems Based on Finite State Machines
by: Zhang, Yaolun, et al.
Published: (2025)
by: Zhang, Yaolun, et al.
Published: (2025)
OET: Optimization-based prompt injection Evaluation Toolkit
by: Pan, Jinsheng, et al.
Published: (2025)
by: Pan, Jinsheng, et al.
Published: (2025)
Explorations in Texture Learning
by: Hoak, Blaine, et al.
Published: (2024)
by: Hoak, Blaine, et al.
Published: (2024)
Robust Graph Contrastive Learning with Information Restoration
by: Zhu, Yulin, et al.
Published: (2023)
by: Zhu, Yulin, et al.
Published: (2023)
Rationality of Learning Algorithms in Repeated Normal-Form Games
by: Bajaj, Shivam, et al.
Published: (2024)
by: Bajaj, Shivam, et al.
Published: (2024)
Protecting Language Models Against Unauthorized Distillation through Trace Rewriting
by: Ma, Xinhang, et al.
Published: (2026)
by: Ma, Xinhang, et al.
Published: (2026)
Learned Neighbor Trust for Collaborative Deployment in Model-Agnostic Decentralized Learning
by: Lanier, Michael, et al.
Published: (2026)
by: Lanier, Michael, et al.
Published: (2026)
To Give or Not to Give? The Impacts of Strategically Withheld Recourse
by: Chen, Yatong, et al.
Published: (2025)
by: Chen, Yatong, et al.
Published: (2025)
Linear Social Choice with Few Queries: A Moment-Based Approach
by: Ge, Luise, et al.
Published: (2026)
by: Ge, Luise, et al.
Published: (2026)
Conformal Reachability for Safe Control in Unknown Environments
by: Ma, Xinhang, et al.
Published: (2026)
by: Ma, Xinhang, et al.
Published: (2026)
Adversarial Machine Unlearning
by: Di, Zonglin, et al.
Published: (2024)
by: Di, Zonglin, et al.
Published: (2024)
Attacks on Node Attributes in Graph Neural Networks
by: Xu, Ying, et al.
Published: (2024)
by: Xu, Ying, et al.
Published: (2024)
Optimal policy for control of epidemics with constrained time intervals and region-based interactions
by: Li, Xia, et al.
Published: (2024)
by: Li, Xia, et al.
Published: (2024)
Similar Items
-
AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
by: Liu, Xiaogeng, et al.
Published: (2025) -
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
by: Liu, Xiaogeng, et al.
Published: (2023) -
RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process
by: Wang, Peiran, et al.
Published: (2024) -
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
by: Wu, Fangzhou, et al.
Published: (2024) -
Verified Safe Reinforcement Learning for Neural Network Dynamic Models
by: Wu, Junlin, et al.
Published: (2024)