Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Kai, Aggarwal, Abhinav, Khodabandeh, Mehran, Zhang, David, Hsin, Eric, Chen, Li, Jain, Ankit, Fredrikson, Matt, Bharadwaj, Akash |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
di: Liu, Yanjiang, et al.
Pubblicazione: (2025)
di: Liu, Yanjiang, et al.
Pubblicazione: (2025)
ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming
di: Béjar, Mario Rodríguez, et al.
Pubblicazione: (2026)
di: Béjar, Mario Rodríguez, et al.
Pubblicazione: (2026)
A Red Teaming Framework for Evaluating Robustness of AI-enabled Security Orchestration, Automation, and Response Systems
di: Shaikh, Ayan Javeed, et al.
Pubblicazione: (2026)
di: Shaikh, Ayan Javeed, et al.
Pubblicazione: (2026)
When the Same Coefficients Reach Different Places: Asymmetric Realizability in Transplanting Tokenizers across Large Language Models
di: Liu, Xiaoze, et al.
Pubblicazione: (2025)
di: Liu, Xiaoze, et al.
Pubblicazione: (2025)
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
di: You, Wenhao, et al.
Pubblicazione: (2025)
di: You, Wenhao, et al.
Pubblicazione: (2025)
TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs
di: Shen, Qingchao, et al.
Pubblicazione: (2026)
di: Shen, Qingchao, et al.
Pubblicazione: (2026)
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts
di: Liu, Yi, et al.
Pubblicazione: (2024)
di: Liu, Yi, et al.
Pubblicazione: (2024)
SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
di: Duan, Zenghao, et al.
Pubblicazione: (2026)
di: Duan, Zenghao, et al.
Pubblicazione: (2026)
SecCodePRM: A Process Reward Model for Code Security
di: Yu, Weichen, et al.
Pubblicazione: (2026)
di: Yu, Weichen, et al.
Pubblicazione: (2026)
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models
di: Hu, Xiaomeng, et al.
Pubblicazione: (2024)
di: Hu, Xiaomeng, et al.
Pubblicazione: (2024)
Resource Consumption Red-Teaming for Large Vision-Language Models
di: Gao, Haoran, et al.
Pubblicazione: (2025)
di: Gao, Haoran, et al.
Pubblicazione: (2025)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
di: Pathade, Chetan
Pubblicazione: (2025)
di: Pathade, Chetan
Pubblicazione: (2025)
Red Teaming Large Reasoning Models
di: Chen, Jiawei, et al.
Pubblicazione: (2025)
di: Chen, Jiawei, et al.
Pubblicazione: (2025)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
di: Chao, Patrick, et al.
Pubblicazione: (2024)
di: Chao, Patrick, et al.
Pubblicazione: (2024)
JULI: Jailbreak Large Language Models by Self-Introspection
di: Wang, Jesson, et al.
Pubblicazione: (2025)
di: Wang, Jesson, et al.
Pubblicazione: (2025)
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
di: Wei, Zhang, et al.
Pubblicazione: (2025)
di: Wei, Zhang, et al.
Pubblicazione: (2025)
A Mixture of Linear Corrections Generates Secure Code
di: Yu, Weichen, et al.
Pubblicazione: (2025)
di: Yu, Weichen, et al.
Pubblicazione: (2025)
VeriSplit: Secure and Practical Offloading of Machine Learning Inferences across IoT Devices
di: Zhang, Han, et al.
Pubblicazione: (2024)
di: Zhang, Han, et al.
Pubblicazione: (2024)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
di: Wang, Haoran, et al.
Pubblicazione: (2023)
di: Wang, Haoran, et al.
Pubblicazione: (2023)
Red Teaming Methodology for Design Obfuscation
di: Liu, Yuntao, et al.
Pubblicazione: (2025)
di: Liu, Yuntao, et al.
Pubblicazione: (2025)
Autonomous Adversary: Red-Teaming in the age of LLM
di: Mamun, Mohammad, et al.
Pubblicazione: (2026)
di: Mamun, Mohammad, et al.
Pubblicazione: (2026)
PrivCode: When Code Generation Meets Differential Privacy
di: Liu, Zheng, et al.
Pubblicazione: (2025)
di: Liu, Zheng, et al.
Pubblicazione: (2025)
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
di: Wang, Zilong, et al.
Pubblicazione: (2025)
di: Wang, Zilong, et al.
Pubblicazione: (2025)
When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models
di: Ou, Haoran, et al.
Pubblicazione: (2025)
di: Ou, Haoran, et al.
Pubblicazione: (2025)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
di: Xu, Huiyu, et al.
Pubblicazione: (2024)
di: Xu, Huiyu, et al.
Pubblicazione: (2024)
Red Teaming AI Red Teaming
di: Majumdar, Subhabrata, et al.
Pubblicazione: (2025)
di: Majumdar, Subhabrata, et al.
Pubblicazione: (2025)
SQL Injection Jailbreak: A Structural Disaster of Large Language Models
di: Zhao, Jiawei, et al.
Pubblicazione: (2024)
di: Zhao, Jiawei, et al.
Pubblicazione: (2024)
Large Language Lobotomy: Jailbreaking Mixture-of-Experts via Expert Silencing
di: Lintelo, Jona te, et al.
Pubblicazione: (2026)
di: Lintelo, Jona te, et al.
Pubblicazione: (2026)
Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
di: Yin, Ziyi, et al.
Pubblicazione: (2025)
di: Yin, Ziyi, et al.
Pubblicazione: (2025)
Incalmo: An Autonomous LLM-assisted System for Red Teaming Multi-Host Networks
di: Singer, Brian, et al.
Pubblicazione: (2025)
di: Singer, Brian, et al.
Pubblicazione: (2025)
RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking
di: Jiang, Yifan, et al.
Pubblicazione: (2024)
di: Jiang, Yifan, et al.
Pubblicazione: (2024)
CORVUS: Red-Teaming Hallucination Detectors via Internal Signal Camouflage in Large Language Models
di: Min, Nay Myat, et al.
Pubblicazione: (2026)
di: Min, Nay Myat, et al.
Pubblicazione: (2026)
JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation
di: Zhang, Shenyi, et al.
Pubblicazione: (2025)
di: Zhang, Shenyi, et al.
Pubblicazione: (2025)
LAAF: Logic-layer Automated Attack Framework A Systematic Red-Teaming Methodology for LPCI Vulnerabilities in Agentic Large Language Model Systems
di: Atta, Hammad, et al.
Pubblicazione: (2026)
di: Atta, Hammad, et al.
Pubblicazione: (2026)
MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
di: Deng, Gelei, et al.
Pubblicazione: (2023)
di: Deng, Gelei, et al.
Pubblicazione: (2023)
SoK: Robustness in Large Language Models against Jailbreak Attacks
di: Xu, Feiyue, et al.
Pubblicazione: (2026)
di: Xu, Feiyue, et al.
Pubblicazione: (2026)
Emoji-Based Jailbreaking of Large Language Models
di: Gopinadh, M P V S, et al.
Pubblicazione: (2026)
di: Gopinadh, M P V S, et al.
Pubblicazione: (2026)
AJAR: Adaptive Jailbreak Architecture for Red-teaming
di: Dou, Yipu, et al.
Pubblicazione: (2026)
di: Dou, Yipu, et al.
Pubblicazione: (2026)
Automated Progressive Red Teaming
di: Jiang, Bojian, et al.
Pubblicazione: (2024)
di: Jiang, Bojian, et al.
Pubblicazione: (2024)
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
di: Sharma, Mrinank, et al.
Pubblicazione: (2025)
di: Sharma, Mrinank, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
di: Liu, Yanjiang, et al.
Pubblicazione: (2025) -
ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming
di: Béjar, Mario Rodríguez, et al.
Pubblicazione: (2026) -
A Red Teaming Framework for Evaluating Robustness of AI-enabled Security Orchestration, Automation, and Response Systems
di: Shaikh, Ayan Javeed, et al.
Pubblicazione: (2026) -
When the Same Coefficients Reach Different Places: Asymmetric Realizability in Transplanting Tokenizers across Large Language Models
di: Liu, Xiaoze, et al.
Pubblicazione: (2025) -
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
di: You, Wenhao, et al.
Pubblicazione: (2025)