Unilogit: Robust Machine Unlearning for LLMs Using Uniform-Target Self-Distillation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Vasilev, Stefan, Herold, Christian, Liao, Baohao, Hashemi, Seyyed Hadi, Khadivi, Shahram, Monz, Christof
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916728961761280
author Vasilev, Stefan
Herold, Christian
Liao, Baohao
Hashemi, Seyyed Hadi
Khadivi, Shahram
Monz, Christof
author_facet Vasilev, Stefan
Herold, Christian
Liao, Baohao
Hashemi, Seyyed Hadi
Khadivi, Shahram
Monz, Christof
contents This paper introduces Unilogit, a novel self-distillation method for machine unlearning in Large Language Models. Unilogit addresses the challenge of selectively forgetting specific information while maintaining overall model utility, a critical task in compliance with data privacy regulations like GDPR. Unlike prior methods that rely on static hyperparameters or starting model outputs, Unilogit dynamically adjusts target logits to achieve a uniform probability for the target token, leveraging the current model's outputs for more accurate self-distillation targets. This approach not only eliminates the need for additional hyperparameters but also enhances the model's ability to approximate the golden targets. Extensive experiments on public benchmarks and an in-house e-commerce dataset demonstrate Unilogit's superior performance in balancing forget and retain objectives, outperforming state-of-the-art methods such as NPO and UnDIAL. Our analysis further reveals Unilogit's robustness across various scenarios, highlighting its practical applicability and effectiveness in achieving efficacious machine unlearning.
format Preprint
id arxiv_https___arxiv_org_abs_2505_06027
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unilogit: Robust Machine Unlearning for LLMs Using Uniform-Target Self-Distillation
Vasilev, Stefan
Herold, Christian
Liao, Baohao
Hashemi, Seyyed Hadi
Khadivi, Shahram
Monz, Christof
Computation and Language
Machine Learning
68T50
I.2.7
This paper introduces Unilogit, a novel self-distillation method for machine unlearning in Large Language Models. Unilogit addresses the challenge of selectively forgetting specific information while maintaining overall model utility, a critical task in compliance with data privacy regulations like GDPR. Unlike prior methods that rely on static hyperparameters or starting model outputs, Unilogit dynamically adjusts target logits to achieve a uniform probability for the target token, leveraging the current model's outputs for more accurate self-distillation targets. This approach not only eliminates the need for additional hyperparameters but also enhances the model's ability to approximate the golden targets. Extensive experiments on public benchmarks and an in-house e-commerce dataset demonstrate Unilogit's superior performance in balancing forget and retain objectives, outperforming state-of-the-art methods such as NPO and UnDIAL. Our analysis further reveals Unilogit's robustness across various scenarios, highlighting its practical applicability and effectiveness in achieving efficacious machine unlearning.
title Unilogit: Robust Machine Unlearning for LLMs Using Uniform-Target Self-Distillation
topic Computation and Language
Machine Learning
68T50
I.2.7
url https://arxiv.org/abs/2505.06027