Learning-Time Encoding Shapes Unlearning in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Ruihan, Garov, Konstantin, Chaudhuri, Kamalika
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908411576188928
author Wu, Ruihan
Garov, Konstantin
Chaudhuri, Kamalika
author_facet Wu, Ruihan
Garov, Konstantin
Chaudhuri, Kamalika
contents As large language models (LLMs) are increasingly deployed in the real world, the ability to ``unlearn'', or remove specific pieces of knowledge post hoc, has become essential for a variety of reasons ranging from privacy regulations to correcting outdated or harmful content. Prior work has proposed unlearning benchmarks and algorithms, and has typically assumed that the training process and the target model are fixed. In this work, we empirically investigate how learning-time choices in knowledge encoding impact the effectiveness of unlearning factual knowledge. Our experiments reveal two key findings: (1) learning with paraphrased descriptions improves unlearning performance and (2) unlearning individual piece of knowledge from a chunk of text is challenging. Our results suggest that learning-time knowledge encoding may play a central role in enabling reliable post-hoc unlearning.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15076
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning-Time Encoding Shapes Unlearning in LLMs
Wu, Ruihan
Garov, Konstantin
Chaudhuri, Kamalika
Computation and Language
Machine Learning
As large language models (LLMs) are increasingly deployed in the real world, the ability to ``unlearn'', or remove specific pieces of knowledge post hoc, has become essential for a variety of reasons ranging from privacy regulations to correcting outdated or harmful content. Prior work has proposed unlearning benchmarks and algorithms, and has typically assumed that the training process and the target model are fixed. In this work, we empirically investigate how learning-time choices in knowledge encoding impact the effectiveness of unlearning factual knowledge. Our experiments reveal two key findings: (1) learning with paraphrased descriptions improves unlearning performance and (2) unlearning individual piece of knowledge from a chunk of text is challenging. Our results suggest that learning-time knowledge encoding may play a central role in enabling reliable post-hoc unlearning.
title Learning-Time Encoding Shapes Unlearning in LLMs
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.15076