SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhuang, Haomin, Zhang, Yihua, Guo, Kehan, Jia, Jinghan, Liu, Gaowen, Liu, Sijia, Zhang, Xiangliang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909666156478464
author Zhuang, Haomin
Zhang, Yihua
Guo, Kehan
Jia, Jinghan
Liu, Gaowen
Liu, Sijia
Zhang, Xiangliang
author_facet Zhuang, Haomin
Zhang, Yihua
Guo, Kehan
Jia, Jinghan
Liu, Gaowen
Liu, Sijia
Zhang, Xiangliang
contents Recent advancements in LLMs unlearning have shown remarkable success in removing unwanted data-model influences while preserving the model's utility for legitimate knowledge. Despite these strides, sparse Mixture-of-Experts (MoE) LLMs--a key subset of the LLM family--have remained unexplored in the context of unlearning. As MoE LLMs are celebrated for their exceptional performance, we ask:How can unlearning be performed effectively and efficiently on MoE LLMs? Our pilot study shows that the dynamic routing nature of MoE LLMs introduces unique challenges, leading to excessive forgetting, uncontrolled knowledge erasure and substantial utility drops when existing unlearning methods are applied. To address this, we propose a novel Selected-Expert Unlearning Framework (SEUF). Through expert attribution, unlearning is concentrated on the most actively engaged experts for the specified knowledge. Concurrently, an anchor loss is applied to the router to stabilize the active state of this targeted expert, ensuring focused and controlled unlearning. SEUF is compatible with various standard unlearning algorithms. Extensive experiments demonstrate that SEUF enhances both forget quality up to 5% and model utility by 35% on MoE LLMs across various benchmarks and LLM architectures (compared to standard unlearning algorithms), while only unlearning 0.06% of the model parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18797
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
Zhuang, Haomin
Zhang, Yihua
Guo, Kehan
Jia, Jinghan
Liu, Gaowen
Liu, Sijia
Zhang, Xiangliang
Machine Learning
Artificial Intelligence
Computation and Language
Recent advancements in LLMs unlearning have shown remarkable success in removing unwanted data-model influences while preserving the model's utility for legitimate knowledge. Despite these strides, sparse Mixture-of-Experts (MoE) LLMs--a key subset of the LLM family--have remained unexplored in the context of unlearning. As MoE LLMs are celebrated for their exceptional performance, we ask:How can unlearning be performed effectively and efficiently on MoE LLMs? Our pilot study shows that the dynamic routing nature of MoE LLMs introduces unique challenges, leading to excessive forgetting, uncontrolled knowledge erasure and substantial utility drops when existing unlearning methods are applied. To address this, we propose a novel Selected-Expert Unlearning Framework (SEUF). Through expert attribution, unlearning is concentrated on the most actively engaged experts for the specified knowledge. Concurrently, an anchor loss is applied to the router to stabilize the active state of this targeted expert, ensuring focused and controlled unlearning. SEUF is compatible with various standard unlearning algorithms. Extensive experiments demonstrate that SEUF enhances both forget quality up to 5% and model utility by 35% on MoE LLMs across various benchmarks and LLM architectures (compared to standard unlearning algorithms), while only unlearning 0.06% of the model parameters.
title SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2411.18797