UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shumailov, Ilia, Hayes, Jamie, Triantafillou, Eleni, Ortiz-Jimenez, Guillermo, Papernot, Nicolas, Jagielski, Matthew, Yona, Itay, Howard, Heidi, Bagdasaryan, Eugene
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929403785641984
author Shumailov, Ilia
Hayes, Jamie
Triantafillou, Eleni
Ortiz-Jimenez, Guillermo
Papernot, Nicolas
Jagielski, Matthew
Yona, Itay
Howard, Heidi
Bagdasaryan, Eugene
author_facet Shumailov, Ilia
Hayes, Jamie
Triantafillou, Eleni
Ortiz-Jimenez, Guillermo
Papernot, Nicolas
Jagielski, Matthew
Yona, Itay
Howard, Heidi
Bagdasaryan, Eugene
contents Exact unlearning was first introduced as a privacy mechanism that allowed a user to retract their data from machine learning models on request. Shortly after, inexact schemes were proposed to mitigate the impractical costs associated with exact unlearning. More recently unlearning is often discussed as an approach for removal of impermissible knowledge i.e. knowledge that the model should not possess such as unlicensed copyrighted, inaccurate, or malicious information. The promise is that if the model does not have a certain malicious capability, then it cannot be used for the associated malicious purpose. In this paper we revisit the paradigm in which unlearning is used for in Large Language Models (LLMs) and highlight an underlying inconsistency arising from in-context learning. Unlearning can be an effective control mechanism for the training phase, yet it does not prevent the model from performing an impermissible act during inference. We introduce a concept of ununlearning, where unlearned knowledge gets reintroduced in-context, effectively rendering the model capable of behaving as if it knows the forgotten knowledge. As a result, we argue that content filtering for impermissible knowledge will be required and even exact unlearning schemes are not enough for effective content regulation. We discuss feasibility of ununlearning for modern LLMs and examine broader implications.
format Preprint
id arxiv_https___arxiv_org_abs_2407_00106
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
Shumailov, Ilia
Hayes, Jamie
Triantafillou, Eleni
Ortiz-Jimenez, Guillermo
Papernot, Nicolas
Jagielski, Matthew
Yona, Itay
Howard, Heidi
Bagdasaryan, Eugene
Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
Exact unlearning was first introduced as a privacy mechanism that allowed a user to retract their data from machine learning models on request. Shortly after, inexact schemes were proposed to mitigate the impractical costs associated with exact unlearning. More recently unlearning is often discussed as an approach for removal of impermissible knowledge i.e. knowledge that the model should not possess such as unlicensed copyrighted, inaccurate, or malicious information. The promise is that if the model does not have a certain malicious capability, then it cannot be used for the associated malicious purpose. In this paper we revisit the paradigm in which unlearning is used for in Large Language Models (LLMs) and highlight an underlying inconsistency arising from in-context learning. Unlearning can be an effective control mechanism for the training phase, yet it does not prevent the model from performing an impermissible act during inference. We introduce a concept of ununlearning, where unlearned knowledge gets reintroduced in-context, effectively rendering the model capable of behaving as if it knows the forgotten knowledge. As a result, we argue that content filtering for impermissible knowledge will be required and even exact unlearning schemes are not enough for effective content regulation. We discuss feasibility of ununlearning for modern LLMs and examine broader implications.
title UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
topic Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2407.00106