Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910807061692416 |
|---|---|
| author | Sharma, Mrinank Tong, Meg Mu, Jesse Wei, Jerry Kruthoff, Jorrit Goodfriend, Scott Ong, Euan Peng, Alwin Agarwal, Raj Anil, Cem Askell, Amanda Bailey, Nathan Benton, Joe Bluemke, Emma Bowman, Samuel R. Christiansen, Eric Cunningham, Hoagy Dau, Andy Gopal, Anjali Gilson, Rob Graham, Logan Howard, Logan Kalra, Nimit Lee, Taesung Lin, Kevin Lofgren, Peter Mosconi, Francesco O'Hara, Clare Olsson, Catherine Petrini, Linda Rajani, Samir Saxena, Nikhil Silverstein, Alex Singh, Tanya Sumers, Theodore Tang, Leonard Troy, Kevin K. Weisser, Constantin Zhong, Ruiqi Zhou, Giulio Leike, Jan Kaplan, Jared Perez, Ethan |
| author_facet | Sharma, Mrinank Tong, Meg Mu, Jesse Wei, Jerry Kruthoff, Jorrit Goodfriend, Scott Ong, Euan Peng, Alwin Agarwal, Raj Anil, Cem Askell, Amanda Bailey, Nathan Benton, Joe Bluemke, Emma Bowman, Samuel R. Christiansen, Eric Cunningham, Hoagy Dau, Andy Gopal, Anjali Gilson, Rob Graham, Logan Howard, Logan Kalra, Nimit Lee, Taesung Lin, Kevin Lofgren, Peter Mosconi, Francesco O'Hara, Clare Olsson, Catherine Petrini, Linda Rajani, Samir Saxena, Nikhil Silverstein, Alex Singh, Tanya Sumers, Theodore Tang, Leonard Troy, Kevin K. Weisser, Constantin Zhong, Ruiqi Zhou, Giulio Leike, Jan Kaplan, Jared Perez, Ethan |
| contents | Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes that require many model interactions, like manufacturing illegal substances at scale. To defend against these attacks, we introduce Constitutional Classifiers: safeguards trained on synthetic data, generated by prompting LLMs with natural language rules (i.e., a constitution) specifying permitted and restricted content. In over 3,000 estimated hours of red teaming, no red teamer found a universal jailbreak that could extract information from an early classifier-guarded LLM at a similar level of detail to an unguarded model across most target queries. On automated evaluations, enhanced classifiers demonstrated robust defense against held-out domain-specific jailbreaks. These classifiers also maintain deployment viability, with an absolute 0.38% increase in production-traffic refusals and a 23.7% inference overhead. Our work demonstrates that defending against universal jailbreaks while maintaining practical deployment viability is tractable. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_18837 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming Sharma, Mrinank Tong, Meg Mu, Jesse Wei, Jerry Kruthoff, Jorrit Goodfriend, Scott Ong, Euan Peng, Alwin Agarwal, Raj Anil, Cem Askell, Amanda Bailey, Nathan Benton, Joe Bluemke, Emma Bowman, Samuel R. Christiansen, Eric Cunningham, Hoagy Dau, Andy Gopal, Anjali Gilson, Rob Graham, Logan Howard, Logan Kalra, Nimit Lee, Taesung Lin, Kevin Lofgren, Peter Mosconi, Francesco O'Hara, Clare Olsson, Catherine Petrini, Linda Rajani, Samir Saxena, Nikhil Silverstein, Alex Singh, Tanya Sumers, Theodore Tang, Leonard Troy, Kevin K. Weisser, Constantin Zhong, Ruiqi Zhou, Giulio Leike, Jan Kaplan, Jared Perez, Ethan Computation and Language Artificial Intelligence Cryptography and Security Machine Learning Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes that require many model interactions, like manufacturing illegal substances at scale. To defend against these attacks, we introduce Constitutional Classifiers: safeguards trained on synthetic data, generated by prompting LLMs with natural language rules (i.e., a constitution) specifying permitted and restricted content. In over 3,000 estimated hours of red teaming, no red teamer found a universal jailbreak that could extract information from an early classifier-guarded LLM at a similar level of detail to an unguarded model across most target queries. On automated evaluations, enhanced classifiers demonstrated robust defense against held-out domain-specific jailbreaks. These classifiers also maintain deployment viability, with an absolute 0.38% increase in production-traffic refusals and a 23.7% inference overhead. Our work demonstrates that defending against universal jailbreaks while maintaining practical deployment viability is tractable. |
| title | Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming |
| topic | Computation and Language Artificial Intelligence Cryptography and Security Machine Learning |
| url | https://arxiv.org/abs/2501.18837 |