Improving Alignment and Robustness with Circuit Breakers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zou, Andy, Phan, Long, Wang, Justin, Duenas, Derek, Lin, Maxwell, Andriushchenko, Maksym, Wang, Rowan, Kolter, Zico, Fredrikson, Matt, Hendrycks, Dan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929418767695872
author Zou, Andy
Phan, Long
Wang, Justin
Duenas, Derek
Lin, Maxwell
Andriushchenko, Maksym
Wang, Rowan
Kolter, Zico
Fredrikson, Matt
Hendrycks, Dan
author_facet Zou, Andy
Phan, Long
Wang, Justin
Duenas, Derek
Lin, Maxwell
Andriushchenko, Maksym
Wang, Rowan
Kolter, Zico
Fredrikson, Matt
Hendrycks, Dan
contents AI systems can take harmful actions and are highly vulnerable to adversarial attacks. We present an approach, inspired by recent advances in representation engineering, that interrupts the models as they respond with harmful outputs with "circuit breakers." Existing techniques aimed at improving alignment, such as refusal training, are often bypassed. Techniques such as adversarial training try to plug these holes by countering specific attacks. As an alternative to refusal training and adversarial training, circuit-breaking directly controls the representations that are responsible for harmful outputs in the first place. Our technique can be applied to both text-only and multimodal language models to prevent the generation of harmful outputs without sacrificing utility -- even in the presence of powerful unseen attacks. Notably, while adversarial robustness in standalone image recognition remains an open challenge, circuit breakers allow the larger multimodal system to reliably withstand image "hijacks" that aim to produce harmful content. Finally, we extend our approach to AI agents, demonstrating considerable reductions in the rate of harmful actions when they are under attack. Our approach represents a significant step forward in the development of reliable safeguards to harmful behavior and adversarial attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04313
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Alignment and Robustness with Circuit Breakers
Zou, Andy
Phan, Long
Wang, Justin
Duenas, Derek
Lin, Maxwell
Andriushchenko, Maksym
Wang, Rowan
Kolter, Zico
Fredrikson, Matt
Hendrycks, Dan
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Computers and Society
AI systems can take harmful actions and are highly vulnerable to adversarial attacks. We present an approach, inspired by recent advances in representation engineering, that interrupts the models as they respond with harmful outputs with "circuit breakers." Existing techniques aimed at improving alignment, such as refusal training, are often bypassed. Techniques such as adversarial training try to plug these holes by countering specific attacks. As an alternative to refusal training and adversarial training, circuit-breaking directly controls the representations that are responsible for harmful outputs in the first place. Our technique can be applied to both text-only and multimodal language models to prevent the generation of harmful outputs without sacrificing utility -- even in the presence of powerful unseen attacks. Notably, while adversarial robustness in standalone image recognition remains an open challenge, circuit breakers allow the larger multimodal system to reliably withstand image "hijacks" that aim to produce harmful content. Finally, we extend our approach to AI agents, demonstrating considerable reductions in the rate of harmful actions when they are under attack. Our approach represents a significant step forward in the development of reliable safeguards to harmful behavior and adversarial attacks.
title Improving Alignment and Robustness with Circuit Breakers
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Computers and Society
url https://arxiv.org/abs/2406.04313