UCD: Unlearning in LLMs via Contrastive Decoding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Suriyakumar, Vinith M., Sekhari, Ayush, Wilson, Ashia
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908407717429248
author Suriyakumar, Vinith M.
Sekhari, Ayush
Wilson, Ashia
author_facet Suriyakumar, Vinith M.
Sekhari, Ayush
Wilson, Ashia
contents Machine unlearning aims to remove specific information, e.g. sensitive or undesirable content, from large language models (LLMs) while preserving overall performance. We propose an inference-time unlearning algorithm that uses contrastive decoding, leveraging two auxiliary smaller models, one trained without the forget set and one trained with it, to guide the outputs of the original model using their difference during inference. Our strategy substantially improves the tradeoff between unlearning effectiveness and model utility. We evaluate our approach on two unlearning benchmarks, TOFU and MUSE. Results show notable gains in both forget quality and retained performance in comparison to prior approaches, suggesting that incorporating contrastive decoding can offer an efficient, practical avenue for unlearning concepts in large-scale models.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12097
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UCD: Unlearning in LLMs via Contrastive Decoding
Suriyakumar, Vinith M.
Sekhari, Ayush
Wilson, Ashia
Computation and Language
Cryptography and Security
Machine Learning
Machine unlearning aims to remove specific information, e.g. sensitive or undesirable content, from large language models (LLMs) while preserving overall performance. We propose an inference-time unlearning algorithm that uses contrastive decoding, leveraging two auxiliary smaller models, one trained without the forget set and one trained with it, to guide the outputs of the original model using their difference during inference. Our strategy substantially improves the tradeoff between unlearning effectiveness and model utility. We evaluate our approach on two unlearning benchmarks, TOFU and MUSE. Results show notable gains in both forget quality and retained performance in comparison to prior approaches, suggesting that incorporating contrastive decoding can offer an efficient, practical avenue for unlearning concepts in large-scale models.
title UCD: Unlearning in LLMs via Contrastive Decoding
topic Computation and Language
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2506.12097