Reliable Unlearning Harmful Information in LLMs with Metamorphosis Representation Projection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Chengcan, Wei, Zeming, Chen, Huanran, Dong, Yinpeng, Sun, Meng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908497134747648
author Wu, Chengcan
Wei, Zeming
Chen, Huanran
Dong, Yinpeng
Sun, Meng
author_facet Wu, Chengcan
Wei, Zeming
Chen, Huanran
Dong, Yinpeng
Sun, Meng
contents While Large Language Models (LLMs) have demonstrated impressive performance in various domains and tasks, concerns about their safety are becoming increasingly severe. In particular, since models may store unsafe knowledge internally, machine unlearning has emerged as a representative paradigm to ensure model safety. Existing approaches employ various training techniques, such as gradient ascent and negative preference optimization, in attempts to eliminate the influence of undesired data on target models. However, these methods merely suppress the activation of undesired data through parametric training without completely eradicating its informational traces within the model. This fundamental limitation makes it difficult to achieve effective continuous unlearning, rendering these methods vulnerable to relearning attacks. To overcome these challenges, we propose a Metamorphosis Representation Projection (MRP) approach that pioneers the application of irreversible projection properties to machine unlearning. By implementing projective transformations in the hidden state space of specific network layers, our method effectively eliminates harmful information while preserving useful knowledge. Experimental results demonstrate that our approach enables effective continuous unlearning and successfully defends against relearning attacks, achieving state-of-the-art performance in unlearning effectiveness while preserving natural performance. Our code is available in https://github.com/ChengcanWu/MRP.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15449
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reliable Unlearning Harmful Information in LLMs with Metamorphosis Representation Projection
Wu, Chengcan
Wei, Zeming
Chen, Huanran
Dong, Yinpeng
Sun, Meng
Machine Learning
Artificial Intelligence
68T07
I.2.6
While Large Language Models (LLMs) have demonstrated impressive performance in various domains and tasks, concerns about their safety are becoming increasingly severe. In particular, since models may store unsafe knowledge internally, machine unlearning has emerged as a representative paradigm to ensure model safety. Existing approaches employ various training techniques, such as gradient ascent and negative preference optimization, in attempts to eliminate the influence of undesired data on target models. However, these methods merely suppress the activation of undesired data through parametric training without completely eradicating its informational traces within the model. This fundamental limitation makes it difficult to achieve effective continuous unlearning, rendering these methods vulnerable to relearning attacks. To overcome these challenges, we propose a Metamorphosis Representation Projection (MRP) approach that pioneers the application of irreversible projection properties to machine unlearning. By implementing projective transformations in the hidden state space of specific network layers, our method effectively eliminates harmful information while preserving useful knowledge. Experimental results demonstrate that our approach enables effective continuous unlearning and successfully defends against relearning attacks, achieving state-of-the-art performance in unlearning effectiveness while preserving natural performance. Our code is available in https://github.com/ChengcanWu/MRP.
title Reliable Unlearning Harmful Information in LLMs with Metamorphosis Representation Projection
topic Machine Learning
Artificial Intelligence
68T07
I.2.6
url https://arxiv.org/abs/2508.15449