Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Nakyeong, Kang, Taegwan, Choi, Jungkyu, Lee, Honglak, Jung, Kyomin
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916274080055296
author Yang, Nakyeong
Kang, Taegwan
Choi, Jungkyu
Lee, Honglak
Jung, Kyomin
author_facet Yang, Nakyeong
Kang, Taegwan
Choi, Jungkyu
Lee, Honglak
Jung, Kyomin
contents Instruction-following language models often show undesirable biases. These undesirable biases may be accelerated in the real-world usage of language models, where a wide range of instructions is used through zero-shot example prompting. To solve this problem, we first define the bias neuron, which significantly affects biased outputs, and prove its existence empirically. Furthermore, we propose a novel and practical bias mitigation method, CRISPR, to eliminate bias neurons of language models in instruction-following settings. CRISPR automatically determines biased outputs and categorizes neurons that affect the biased outputs as bias neurons using an explainability method. Experimental results demonstrate the effectiveness of our method in mitigating biases under zero-shot instruction-following settings without losing the model's task performance and existing knowledge. The experimental results reveal the generalizability of our method as it shows robustness under various instructions and datasets. Surprisingly, our method can mitigate the bias in language models by eliminating only a few neurons (at least three).
format Preprint
id arxiv_https___arxiv_org_abs_2311_09627
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination
Yang, Nakyeong
Kang, Taegwan
Choi, Jungkyu
Lee, Honglak
Jung, Kyomin
Artificial Intelligence
Computation and Language
Machine Learning
Instruction-following language models often show undesirable biases. These undesirable biases may be accelerated in the real-world usage of language models, where a wide range of instructions is used through zero-shot example prompting. To solve this problem, we first define the bias neuron, which significantly affects biased outputs, and prove its existence empirically. Furthermore, we propose a novel and practical bias mitigation method, CRISPR, to eliminate bias neurons of language models in instruction-following settings. CRISPR automatically determines biased outputs and categorizes neurons that affect the biased outputs as bias neurons using an explainability method. Experimental results demonstrate the effectiveness of our method in mitigating biases under zero-shot instruction-following settings without losing the model's task performance and existing knowledge. The experimental results reveal the generalizability of our method as it shows robustness under various instructions and datasets. Surprisingly, our method can mitigate the bias in language models by eliminating only a few neurons (at least three).
title Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2311.09627