Separate the Wheat from the Chaff: Model Deficiency Unlearning via Parameter-Efficient Module Operation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hu, Xinshuo, Li, Dongfang, Hu, Baotian, Zheng, Zihao, Liu, Zhenyu, Zhang, Min
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911760026435584
author Hu, Xinshuo
Li, Dongfang
Hu, Baotian
Zheng, Zihao
Liu, Zhenyu
Zhang, Min
author_facet Hu, Xinshuo
Li, Dongfang
Hu, Baotian
Zheng, Zihao
Liu, Zhenyu
Zhang, Min
contents Large language models (LLMs) have been widely used in various applications but are known to suffer from issues related to untruthfulness and toxicity. While parameter-efficient modules (PEMs) have demonstrated their effectiveness in equipping models with new skills, leveraging PEMs for deficiency unlearning remains underexplored. In this work, we propose a PEMs operation approach, namely Extraction-before-Subtraction (Ext-Sub), to enhance the truthfulness and detoxification of LLMs through the integration of ``expert'' PEM and ``anti-expert'' PEM. Remarkably, even anti-expert PEM possess valuable capabilities due to their proficiency in generating fabricated content, which necessitates language modeling and logical narrative competence. Rather than merely negating the parameters, our approach involves extracting and eliminating solely the deficiency capability within anti-expert PEM while preserving the general capabilities. To evaluate the effectiveness of our approach in terms of truthfulness and detoxification, we conduct extensive experiments on LLMs, encompassing additional abilities such as language modeling and mathematical reasoning. Our empirical results demonstrate that our approach effectively improves truthfulness and detoxification, while largely preserving the fundamental abilities of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2308_08090
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Separate the Wheat from the Chaff: Model Deficiency Unlearning via Parameter-Efficient Module Operation
Hu, Xinshuo
Li, Dongfang
Hu, Baotian
Zheng, Zihao
Liu, Zhenyu
Zhang, Min
Computation and Language
Large language models (LLMs) have been widely used in various applications but are known to suffer from issues related to untruthfulness and toxicity. While parameter-efficient modules (PEMs) have demonstrated their effectiveness in equipping models with new skills, leveraging PEMs for deficiency unlearning remains underexplored. In this work, we propose a PEMs operation approach, namely Extraction-before-Subtraction (Ext-Sub), to enhance the truthfulness and detoxification of LLMs through the integration of ``expert'' PEM and ``anti-expert'' PEM. Remarkably, even anti-expert PEM possess valuable capabilities due to their proficiency in generating fabricated content, which necessitates language modeling and logical narrative competence. Rather than merely negating the parameters, our approach involves extracting and eliminating solely the deficiency capability within anti-expert PEM while preserving the general capabilities. To evaluate the effectiveness of our approach in terms of truthfulness and detoxification, we conduct extensive experiments on LLMs, encompassing additional abilities such as language modeling and mathematical reasoning. Our empirical results demonstrate that our approach effectively improves truthfulness and detoxification, while largely preserving the fundamental abilities of LLMs.
title Separate the Wheat from the Chaff: Model Deficiency Unlearning via Parameter-Efficient Module Operation
topic Computation and Language
url https://arxiv.org/abs/2308.08090