_version_ 1866910807061692416
author Sharma, Mrinank
Tong, Meg
Mu, Jesse
Wei, Jerry
Kruthoff, Jorrit
Goodfriend, Scott
Ong, Euan
Peng, Alwin
Agarwal, Raj
Anil, Cem
Askell, Amanda
Bailey, Nathan
Benton, Joe
Bluemke, Emma
Bowman, Samuel R.
Christiansen, Eric
Cunningham, Hoagy
Dau, Andy
Gopal, Anjali
Gilson, Rob
Graham, Logan
Howard, Logan
Kalra, Nimit
Lee, Taesung
Lin, Kevin
Lofgren, Peter
Mosconi, Francesco
O'Hara, Clare
Olsson, Catherine
Petrini, Linda
Rajani, Samir
Saxena, Nikhil
Silverstein, Alex
Singh, Tanya
Sumers, Theodore
Tang, Leonard
Troy, Kevin K.
Weisser, Constantin
Zhong, Ruiqi
Zhou, Giulio
Leike, Jan
Kaplan, Jared
Perez, Ethan
author_facet Sharma, Mrinank
Tong, Meg
Mu, Jesse
Wei, Jerry
Kruthoff, Jorrit
Goodfriend, Scott
Ong, Euan
Peng, Alwin
Agarwal, Raj
Anil, Cem
Askell, Amanda
Bailey, Nathan
Benton, Joe
Bluemke, Emma
Bowman, Samuel R.
Christiansen, Eric
Cunningham, Hoagy
Dau, Andy
Gopal, Anjali
Gilson, Rob
Graham, Logan
Howard, Logan
Kalra, Nimit
Lee, Taesung
Lin, Kevin
Lofgren, Peter
Mosconi, Francesco
O'Hara, Clare
Olsson, Catherine
Petrini, Linda
Rajani, Samir
Saxena, Nikhil
Silverstein, Alex
Singh, Tanya
Sumers, Theodore
Tang, Leonard
Troy, Kevin K.
Weisser, Constantin
Zhong, Ruiqi
Zhou, Giulio
Leike, Jan
Kaplan, Jared
Perez, Ethan
contents Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes that require many model interactions, like manufacturing illegal substances at scale. To defend against these attacks, we introduce Constitutional Classifiers: safeguards trained on synthetic data, generated by prompting LLMs with natural language rules (i.e., a constitution) specifying permitted and restricted content. In over 3,000 estimated hours of red teaming, no red teamer found a universal jailbreak that could extract information from an early classifier-guarded LLM at a similar level of detail to an unguarded model across most target queries. On automated evaluations, enhanced classifiers demonstrated robust defense against held-out domain-specific jailbreaks. These classifiers also maintain deployment viability, with an absolute 0.38% increase in production-traffic refusals and a 23.7% inference overhead. Our work demonstrates that defending against universal jailbreaks while maintaining practical deployment viability is tractable.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18837
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Sharma, Mrinank
Tong, Meg
Mu, Jesse
Wei, Jerry
Kruthoff, Jorrit
Goodfriend, Scott
Ong, Euan
Peng, Alwin
Agarwal, Raj
Anil, Cem
Askell, Amanda
Bailey, Nathan
Benton, Joe
Bluemke, Emma
Bowman, Samuel R.
Christiansen, Eric
Cunningham, Hoagy
Dau, Andy
Gopal, Anjali
Gilson, Rob
Graham, Logan
Howard, Logan
Kalra, Nimit
Lee, Taesung
Lin, Kevin
Lofgren, Peter
Mosconi, Francesco
O'Hara, Clare
Olsson, Catherine
Petrini, Linda
Rajani, Samir
Saxena, Nikhil
Silverstein, Alex
Singh, Tanya
Sumers, Theodore
Tang, Leonard
Troy, Kevin K.
Weisser, Constantin
Zhong, Ruiqi
Zhou, Giulio
Leike, Jan
Kaplan, Jared
Perez, Ethan
Computation and Language
Artificial Intelligence
Cryptography and Security
Machine Learning
Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes that require many model interactions, like manufacturing illegal substances at scale. To defend against these attacks, we introduce Constitutional Classifiers: safeguards trained on synthetic data, generated by prompting LLMs with natural language rules (i.e., a constitution) specifying permitted and restricted content. In over 3,000 estimated hours of red teaming, no red teamer found a universal jailbreak that could extract information from an early classifier-guarded LLM at a similar level of detail to an unguarded model across most target queries. On automated evaluations, enhanced classifiers demonstrated robust defense against held-out domain-specific jailbreaks. These classifiers also maintain deployment viability, with an absolute 0.38% increase in production-traffic refusals and a 23.7% inference overhead. Our work demonstrates that defending against universal jailbreaks while maintaining practical deployment viability is tractable.
title Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
topic Computation and Language
Artificial Intelligence
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2501.18837