LLM Surgery: Efficient Knowledge Unlearning and Editing in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Veldanda, Akshaj Kumar, Zhang, Shi-Xiong, Das, Anirban, Chakraborty, Supriyo, Rawls, Stephen, Sahu, Sambit, Naphade, Milind
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916402466652160
author Veldanda, Akshaj Kumar
Zhang, Shi-Xiong
Das, Anirban
Chakraborty, Supriyo
Rawls, Stephen
Sahu, Sambit
Naphade, Milind
author_facet Veldanda, Akshaj Kumar
Zhang, Shi-Xiong
Das, Anirban
Chakraborty, Supriyo
Rawls, Stephen
Sahu, Sambit
Naphade, Milind
contents Large language models (LLMs) have revolutionized various domains, yet their utility comes with significant challenges related to outdated or problematic knowledge embedded during pretraining. This paper addresses the challenge of modifying LLMs to unlearn problematic and outdated information while efficiently integrating new knowledge without retraining from scratch. Here, we propose LLM Surgery, a framework to efficiently modify LLM behaviour by optimizing a three component objective function that: (1) Performs reverse gradient on unlearning dataset (problematic and outdated information), (2) Performs gradient descent on the update dataset (new and updated information), and (3) Minimizes the KL divergence on the retain dataset (small subset of unchanged text), ensuring alignment between pretrained and modified model outputs. Due to the lack of publicly available datasets specifically tailored for our novel task, we compiled a new dataset and an evaluation benchmark. Using Llama2-7B, we demonstrate that LLM Surgery can achieve significant forgetting on the unlearn set, a 20\% increase in accuracy on the update set, and maintain performance on the retain set.
format Preprint
id arxiv_https___arxiv_org_abs_2409_13054
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLM Surgery: Efficient Knowledge Unlearning and Editing in Large Language Models
Veldanda, Akshaj Kumar
Zhang, Shi-Xiong
Das, Anirban
Chakraborty, Supriyo
Rawls, Stephen
Sahu, Sambit
Naphade, Milind
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) have revolutionized various domains, yet their utility comes with significant challenges related to outdated or problematic knowledge embedded during pretraining. This paper addresses the challenge of modifying LLMs to unlearn problematic and outdated information while efficiently integrating new knowledge without retraining from scratch. Here, we propose LLM Surgery, a framework to efficiently modify LLM behaviour by optimizing a three component objective function that: (1) Performs reverse gradient on unlearning dataset (problematic and outdated information), (2) Performs gradient descent on the update dataset (new and updated information), and (3) Minimizes the KL divergence on the retain dataset (small subset of unchanged text), ensuring alignment between pretrained and modified model outputs. Due to the lack of publicly available datasets specifically tailored for our novel task, we compiled a new dataset and an evaluation benchmark. Using Llama2-7B, we demonstrate that LLM Surgery can achieve significant forgetting on the unlearn set, a 20\% increase in accuracy on the update set, and maintain performance on the retain set.
title LLM Surgery: Efficient Knowledge Unlearning and Editing in Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2409.13054