Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Yanjiang, Zhou, Shuhen, Lu, Yaojie, Zhu, Huijia, Wang, Weiqiang, Lin, Hongyu, He, Ben, Han, Xianpei, Sun, Le |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
di: You, Wenhao, et al.
Pubblicazione: (2025)
di: You, Wenhao, et al.
Pubblicazione: (2025)
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts
di: Liu, Yi, et al.
Pubblicazione: (2024)
di: Liu, Yi, et al.
Pubblicazione: (2024)
AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models
di: Gautam, Tanmay, et al.
Pubblicazione: (2026)
di: Gautam, Tanmay, et al.
Pubblicazione: (2026)
FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption
di: Wang, Yanting, et al.
Pubblicazione: (2026)
di: Wang, Yanting, et al.
Pubblicazione: (2026)
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
di: Hu, Kai, et al.
Pubblicazione: (2025)
di: Hu, Kai, et al.
Pubblicazione: (2025)
AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
di: Liu, Xiaogeng, et al.
Pubblicazione: (2025)
di: Liu, Xiaogeng, et al.
Pubblicazione: (2025)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
di: Lu, Lin, et al.
Pubblicazione: (2024)
di: Lu, Lin, et al.
Pubblicazione: (2024)
CORVUS: Red-Teaming Hallucination Detectors via Internal Signal Camouflage in Large Language Models
di: Min, Nay Myat, et al.
Pubblicazione: (2026)
di: Min, Nay Myat, et al.
Pubblicazione: (2026)
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
di: Zhou, Andy, et al.
Pubblicazione: (2025)
di: Zhou, Andy, et al.
Pubblicazione: (2025)
ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming
di: Béjar, Mario Rodríguez, et al.
Pubblicazione: (2026)
di: Béjar, Mario Rodríguez, et al.
Pubblicazione: (2026)
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
di: Liu, Xiaogeng, et al.
Pubblicazione: (2024)
di: Liu, Xiaogeng, et al.
Pubblicazione: (2024)
OpenRT: An Open-Source Red Teaming Framework for Multimodal LLMs
di: Wang, Xin, et al.
Pubblicazione: (2026)
di: Wang, Xin, et al.
Pubblicazione: (2026)
When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models
di: Ou, Haoran, et al.
Pubblicazione: (2025)
di: Ou, Haoran, et al.
Pubblicazione: (2025)
Red-Teaming LLM Multi-Agent Systems via Communication Attacks
di: He, Pengfei, et al.
Pubblicazione: (2025)
di: He, Pengfei, et al.
Pubblicazione: (2025)
Red Teaming Methodology for Design Obfuscation
di: Liu, Yuntao, et al.
Pubblicazione: (2025)
di: Liu, Yuntao, et al.
Pubblicazione: (2025)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
di: Kumar, Anurakt, et al.
Pubblicazione: (2024)
di: Kumar, Anurakt, et al.
Pubblicazione: (2024)
DeepInception: Hypnotize Large Language Model to Be Jailbreaker
di: Li, Xuan, et al.
Pubblicazione: (2023)
di: Li, Xuan, et al.
Pubblicazione: (2023)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
di: Reddy, Aashray, et al.
Pubblicazione: (2025)
di: Reddy, Aashray, et al.
Pubblicazione: (2025)
Distract Large Language Models for Automatic Jailbreak Attack
di: Xiao, Zeguan, et al.
Pubblicazione: (2024)
di: Xiao, Zeguan, et al.
Pubblicazione: (2024)
Resource Consumption Red-Teaming for Large Vision-Language Models
di: Gao, Haoran, et al.
Pubblicazione: (2025)
di: Gao, Haoran, et al.
Pubblicazione: (2025)
AutoFirm: Automatically Identifying Reused Libraries inside IoT Firmware at Large-Scale
di: Chen, YongLe, et al.
Pubblicazione: (2024)
di: Chen, YongLe, et al.
Pubblicazione: (2024)
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
di: Wei, Zhang, et al.
Pubblicazione: (2025)
di: Wei, Zhang, et al.
Pubblicazione: (2025)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
di: Pathade, Chetan
Pubblicazione: (2025)
di: Pathade, Chetan
Pubblicazione: (2025)
Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
di: Yin, Ziyi, et al.
Pubblicazione: (2025)
di: Yin, Ziyi, et al.
Pubblicazione: (2025)
Red Teaming Large Reasoning Models
di: Chen, Jiawei, et al.
Pubblicazione: (2025)
di: Chen, Jiawei, et al.
Pubblicazione: (2025)
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
di: Chen, Shuo, et al.
Pubblicazione: (2024)
di: Chen, Shuo, et al.
Pubblicazione: (2024)
TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs
di: Shen, Qingchao, et al.
Pubblicazione: (2026)
di: Shen, Qingchao, et al.
Pubblicazione: (2026)
Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
di: He, Ping, et al.
Pubblicazione: (2025)
di: He, Ping, et al.
Pubblicazione: (2025)
Autonomous Adversary: Red-Teaming in the age of LLM
di: Mamun, Mohammad, et al.
Pubblicazione: (2026)
di: Mamun, Mohammad, et al.
Pubblicazione: (2026)
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
di: Wang, Zilong, et al.
Pubblicazione: (2025)
di: Wang, Zilong, et al.
Pubblicazione: (2025)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
di: Xu, Huiyu, et al.
Pubblicazione: (2024)
di: Xu, Huiyu, et al.
Pubblicazione: (2024)
Knowledge-Driven Multi-Turn Jailbreaking on Large Language Models
di: Li, Songze, et al.
Pubblicazione: (2026)
di: Li, Songze, et al.
Pubblicazione: (2026)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
di: Chao, Patrick, et al.
Pubblicazione: (2024)
di: Chao, Patrick, et al.
Pubblicazione: (2024)
Red Teaming AI Red Teaming
di: Majumdar, Subhabrata, et al.
Pubblicazione: (2025)
di: Majumdar, Subhabrata, et al.
Pubblicazione: (2025)
SQL Injection Jailbreak: A Structural Disaster of Large Language Models
di: Zhao, Jiawei, et al.
Pubblicazione: (2024)
di: Zhao, Jiawei, et al.
Pubblicazione: (2024)
Large Language Lobotomy: Jailbreaking Mixture-of-Experts via Expert Silencing
di: Lintelo, Jona te, et al.
Pubblicazione: (2026)
di: Lintelo, Jona te, et al.
Pubblicazione: (2026)
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models
di: Hu, Xiaomeng, et al.
Pubblicazione: (2024)
di: Hu, Xiaomeng, et al.
Pubblicazione: (2024)
The Dark Side of Function Calling: Pathways to Jailbreaking Large Language Models
di: Wu, Zihui, et al.
Pubblicazione: (2024)
di: Wu, Zihui, et al.
Pubblicazione: (2024)
SoK: Robustness in Large Language Models against Jailbreak Attacks
di: Xu, Feiyue, et al.
Pubblicazione: (2026)
di: Xu, Feiyue, et al.
Pubblicazione: (2026)
Universal Jailbreak Suffixes Are Strong Attention Hijackers
di: Ben-Tov, Matan, et al.
Pubblicazione: (2025)
di: Ben-Tov, Matan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
di: You, Wenhao, et al.
Pubblicazione: (2025) -
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts
di: Liu, Yi, et al.
Pubblicazione: (2024) -
AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models
di: Gautam, Tanmay, et al.
Pubblicazione: (2026) -
FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption
di: Wang, Yanting, et al.
Pubblicazione: (2026) -
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
di: Hu, Kai, et al.
Pubblicazione: (2025)