Answer When Needed, Forget When Not: Language Models Pretend to Forget via In-Context Knowledge Unlearning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Takashiro, Shota, Kojima, Takeshi, Gambardella, Andrew, Cao, Qi, Iwasawa, Yusuke, Matsuo, Yutaka
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916773376294912
author Takashiro, Shota
Kojima, Takeshi
Gambardella, Andrew
Cao, Qi
Iwasawa, Yusuke
Matsuo, Yutaka
author_facet Takashiro, Shota
Kojima, Takeshi
Gambardella, Andrew
Cao, Qi
Iwasawa, Yusuke
Matsuo, Yutaka
contents As large language models (LLMs) are applied across diverse domains, the ability to selectively unlearn specific information is becoming increasingly essential. For instance, LLMs are expected to selectively provide confidential information to authorized internal users, such as employees or trusted partners, while withholding it from external users, including the general public and unauthorized entities. Therefore, we propose a novel method termed ``in-context knowledge unlearning'', which enables the model to selectively forget information in test-time based on the query context. Our method fine-tunes pre-trained LLMs to enable prompt unlearning of target knowledge within the context, while preserving unrelated information. Experiments on TOFU, AGE and RWKU datasets using Llama2-7B/13B and Mistral-7B models demonstrate that our method achieves up to 95% forget accuracy while retaining 80% of unrelated knowledge, significantly outperforming baselines in both in-domain and out-of-domain scenarios. Further investigation of the model's internal behavior revealed that while fine-tuned LLMs generate correct predictions in the middle layers and preserve them up to the final layer. However, the decision to forget is made only at the last layer, i.e. ``LLMs pretend to forget''. Our findings offer valuable insight into the improvement of the robustness of the unlearning mechanisms in LLMs, laying a foundation for future research in the field.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00382
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Answer When Needed, Forget When Not: Language Models Pretend to Forget via In-Context Knowledge Unlearning
Takashiro, Shota
Kojima, Takeshi
Gambardella, Andrew
Cao, Qi
Iwasawa, Yusuke
Matsuo, Yutaka
Computation and Language
As large language models (LLMs) are applied across diverse domains, the ability to selectively unlearn specific information is becoming increasingly essential. For instance, LLMs are expected to selectively provide confidential information to authorized internal users, such as employees or trusted partners, while withholding it from external users, including the general public and unauthorized entities. Therefore, we propose a novel method termed ``in-context knowledge unlearning'', which enables the model to selectively forget information in test-time based on the query context. Our method fine-tunes pre-trained LLMs to enable prompt unlearning of target knowledge within the context, while preserving unrelated information. Experiments on TOFU, AGE and RWKU datasets using Llama2-7B/13B and Mistral-7B models demonstrate that our method achieves up to 95% forget accuracy while retaining 80% of unrelated knowledge, significantly outperforming baselines in both in-domain and out-of-domain scenarios. Further investigation of the model's internal behavior revealed that while fine-tuned LLMs generate correct predictions in the middle layers and preserve them up to the final layer. However, the decision to forget is made only at the last layer, i.e. ``LLMs pretend to forget''. Our findings offer valuable insight into the improvement of the robustness of the unlearning mechanisms in LLMs, laying a foundation for future research in the field.
title Answer When Needed, Forget When Not: Language Models Pretend to Forget via In-Context Knowledge Unlearning
topic Computation and Language
url https://arxiv.org/abs/2410.00382