An Adversarial Perspective on Machine Unlearning for AI Safety

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Łucki, Jakub, Wei, Boyi, Huang, Yangsibo, Henderson, Peter, Tramèr, Florian, Rando, Javier
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915314495651840
author Łucki, Jakub
Wei, Boyi
Huang, Yangsibo
Henderson, Peter
Tramèr, Florian
Rando, Javier
author_facet Łucki, Jakub
Wei, Boyi
Huang, Yangsibo
Henderson, Peter
Tramèr, Florian
Rando, Javier
contents Large language models are finetuned to refuse questions about hazardous knowledge, but these protections can often be bypassed. Unlearning methods aim at completely removing hazardous capabilities from models and make them inaccessible to adversaries. This work challenges the fundamental differences between unlearning and traditional safety post-training from an adversarial perspective. We demonstrate that existing jailbreak methods, previously reported as ineffective against unlearning, can be successful when applied carefully. Furthermore, we develop a variety of adaptive methods that recover most supposedly unlearned capabilities. For instance, we show that finetuning on 10 unrelated examples or removing specific directions in the activation space can recover most hazardous capabilities for models edited with RMU, a state-of-the-art unlearning method. Our findings challenge the robustness of current unlearning approaches and question their advantages over safety training.
format Preprint
id arxiv_https___arxiv_org_abs_2409_18025
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Adversarial Perspective on Machine Unlearning for AI Safety
Łucki, Jakub
Wei, Boyi
Huang, Yangsibo
Henderson, Peter
Tramèr, Florian
Rando, Javier
Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
Large language models are finetuned to refuse questions about hazardous knowledge, but these protections can often be bypassed. Unlearning methods aim at completely removing hazardous capabilities from models and make them inaccessible to adversaries. This work challenges the fundamental differences between unlearning and traditional safety post-training from an adversarial perspective. We demonstrate that existing jailbreak methods, previously reported as ineffective against unlearning, can be successful when applied carefully. Furthermore, we develop a variety of adaptive methods that recover most supposedly unlearned capabilities. For instance, we show that finetuning on 10 unrelated examples or removing specific directions in the activation space can recover most hazardous capabilities for models edited with RMU, a state-of-the-art unlearning method. Our findings challenge the robustness of current unlearning approaches and question their advantages over safety training.
title An Adversarial Perspective on Machine Unlearning for AI Safety
topic Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2409.18025