AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Reddy, Aashray, Zagula, Andrew, Saban, Nicholas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
por: Reddy, Aashray, et al.
Publicado: (2025)
por: Reddy, Aashray, et al.
Publicado: (2025)
Knowledge-Driven Multi-Turn Jailbreaking on Large Language Models
por: Li, Songze, et al.
Publicado: (2026)
por: Li, Songze, et al.
Publicado: (2026)
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts
por: Liu, Yi, et al.
Publicado: (2024)
por: Liu, Yi, et al.
Publicado: (2024)
RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking
por: Jiang, Yifan, et al.
Publicado: (2024)
por: Jiang, Yifan, et al.
Publicado: (2024)
Defending Jailbreak Prompts via In-Context Adversarial Game
por: Zhou, Yujun, et al.
Publicado: (2024)
por: Zhou, Yujun, et al.
Publicado: (2024)
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
por: Shen, Xinyue, et al.
Publicado: (2023)
por: Shen, Xinyue, et al.
Publicado: (2023)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
por: Chao, Patrick, et al.
Publicado: (2024)
por: Chao, Patrick, et al.
Publicado: (2024)
Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs
por: Lu, Yiyang, et al.
Publicado: (2026)
por: Lu, Yiyang, et al.
Publicado: (2026)
Jailbreaking Large Language Models in Infinitely Many Ways
por: Goldstein, Oliver, et al.
Publicado: (2025)
por: Goldstein, Oliver, et al.
Publicado: (2025)
JULI: Jailbreak Large Language Models by Self-Introspection
por: Wang, Jesson, et al.
Publicado: (2025)
por: Wang, Jesson, et al.
Publicado: (2025)
DeepInception: Hypnotize Large Language Model to Be Jailbreaker
por: Li, Xuan, et al.
Publicado: (2023)
por: Li, Xuan, et al.
Publicado: (2023)
Auto-ART: Structured Literature Synthesis and Automated Adversarial Robustness Testing
por: Talluri, Abhijit
Publicado: (2026)
por: Talluri, Abhijit
Publicado: (2026)
Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models
por: Wang, Xiangwen, et al.
Publicado: (2026)
por: Wang, Xiangwen, et al.
Publicado: (2026)
AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models
por: Liang, Jiacheng, et al.
Publicado: (2025)
por: Liang, Jiacheng, et al.
Publicado: (2025)
Model-Editing-Based Jailbreak against Safety-aligned Large Language Models
por: Li, Yuxi, et al.
Publicado: (2024)
por: Li, Yuxi, et al.
Publicado: (2024)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
por: Zhu, Kaijie, et al.
Publicado: (2023)
por: Zhu, Kaijie, et al.
Publicado: (2023)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
por: Li, Qizhang, et al.
Publicado: (2024)
por: Li, Qizhang, et al.
Publicado: (2024)
Prompt Obfuscation for Large Language Models
por: Pape, David, et al.
Publicado: (2024)
por: Pape, David, et al.
Publicado: (2024)
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
por: Wu, Yuanwei, et al.
Publicado: (2023)
por: Wu, Yuanwei, et al.
Publicado: (2023)
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
por: Peng, Benji, et al.
Publicado: (2024)
por: Peng, Benji, et al.
Publicado: (2024)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
por: Zeng, Yifan, et al.
Publicado: (2024)
por: Zeng, Yifan, et al.
Publicado: (2024)
Persona-Conditioned Adversarial Prompting: Multi-Identity Red-Teaming for Adversarial Discovery and Mitigation
por: Morasso, Cristian, et al.
Publicado: (2026)
por: Morasso, Cristian, et al.
Publicado: (2026)
Adversarial Search Engine Optimization for Large Language Models
por: Nestaas, Fredrik, et al.
Publicado: (2024)
por: Nestaas, Fredrik, et al.
Publicado: (2024)
Jailbreaking with Universal Multi-Prompts
por: Hsu, Yu-Ling, et al.
Publicado: (2025)
por: Hsu, Yu-Ling, et al.
Publicado: (2025)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
por: Saiem, Bijoy Ahmed, et al.
Publicado: (2024)
por: Saiem, Bijoy Ahmed, et al.
Publicado: (2024)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
por: Nasr, Milad, et al.
Publicado: (2025)
por: Nasr, Milad, et al.
Publicado: (2025)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
por: Mo, Yichuan, et al.
Publicado: (2024)
por: Mo, Yichuan, et al.
Publicado: (2024)
ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs
por: Liu, Xu, et al.
Publicado: (2025)
por: Liu, Xu, et al.
Publicado: (2025)
VERA: Variational Inference Framework for Jailbreaking Large Language Models
por: Lochab, Anamika, et al.
Publicado: (2025)
por: Lochab, Anamika, et al.
Publicado: (2025)
Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
por: Jia, Xiaojun, et al.
Publicado: (2024)
por: Jia, Xiaojun, et al.
Publicado: (2024)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
por: Li, Nathaniel, et al.
Publicado: (2024)
por: Li, Nathaniel, et al.
Publicado: (2024)
L-AutoDA: Leveraging Large Language Models for Automated Decision-based Adversarial Attacks
por: Guo, Ping, et al.
Publicado: (2024)
por: Guo, Ping, et al.
Publicado: (2024)
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
por: Xiang, Zhen, et al.
Publicado: (2024)
por: Xiang, Zhen, et al.
Publicado: (2024)
Improved Large Language Model Jailbreak Detection via Pretrained Embeddings
por: Galinkin, Erick, et al.
Publicado: (2024)
por: Galinkin, Erick, et al.
Publicado: (2024)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
por: Carlini, Nicholas, et al.
Publicado: (2025)
por: Carlini, Nicholas, et al.
Publicado: (2025)
AdvSGM: Differentially Private Graph Learning via Adversarial Skip-gram Model
por: Zhang, Sen, et al.
Publicado: (2025)
por: Zhang, Sen, et al.
Publicado: (2025)
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
por: Wahed, Muntasir, et al.
Publicado: (2025)
por: Wahed, Muntasir, et al.
Publicado: (2025)
FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts
por: Zhang, Ziyi, et al.
Publicado: (2025)
por: Zhang, Ziyi, et al.
Publicado: (2025)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
por: Paulus, Anselm, et al.
Publicado: (2024)
por: Paulus, Anselm, et al.
Publicado: (2024)
Backdoored Retrievers for Prompt Injection Attacks on Retrieval Augmented Generation of Large Language Models
por: Clop, Cody, et al.
Publicado: (2024)
por: Clop, Cody, et al.
Publicado: (2024)
Ejemplares similares
-
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
por: Reddy, Aashray, et al.
Publicado: (2025) -
Knowledge-Driven Multi-Turn Jailbreaking on Large Language Models
por: Li, Songze, et al.
Publicado: (2026) -
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts
por: Liu, Yi, et al.
Publicado: (2024) -
RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking
por: Jiang, Yifan, et al.
Publicado: (2024) -
Defending Jailbreak Prompts via In-Context Adversarial Game
por: Zhou, Yujun, et al.
Publicado: (2024)