How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Yushi, Sondej, Filip, Mayne, Harry, Lee, Andrew, Mahdi, Adam
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913884106915840
author Yang, Yushi
Sondej, Filip
Mayne, Harry
Lee, Andrew
Mahdi, Adam
author_facet Yang, Yushi
Sondej, Filip
Mayne, Harry
Lee, Andrew
Mahdi, Adam
contents Safety fine-tuning algorithms reduce harmful outputs in language models, yet their mechanisms remain under-explored. Direct Preference Optimization (DPO) is a popular choice of algorithm, but prior explanations, attributing its effects solely to dampened toxic neurons in the MLP layers, are incomplete. In this study, we analyse four language models (Llama-3.1-8B, Gemma-2-2B, Mistral-7B, GPT-2-Medium) and show that toxic neurons only account for 2.5% to 24% of DPO's effects across models. Instead, DPO balances distributed activation shifts across all MLP neurons to create a net toxicity reduction. We attribute this reduction to four neuron groups, two aligned with reducing toxicity and two promoting anti-toxicity, whose combined effects replicate DPO across models. To further validate this understanding, we develop an activation editing method mimicking DPO through distributed shifts along a toxicity representation. This method outperforms DPO in reducing toxicity while preserving perplexity, without requiring any weight updates. Our work provides a mechanistic understanding of DPO and introduces an efficient, tuning-free alternative for safety fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2411_06424
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
Yang, Yushi
Sondej, Filip
Mayne, Harry
Lee, Andrew
Mahdi, Adam
Machine Learning
Computation and Language
Safety fine-tuning algorithms reduce harmful outputs in language models, yet their mechanisms remain under-explored. Direct Preference Optimization (DPO) is a popular choice of algorithm, but prior explanations, attributing its effects solely to dampened toxic neurons in the MLP layers, are incomplete. In this study, we analyse four language models (Llama-3.1-8B, Gemma-2-2B, Mistral-7B, GPT-2-Medium) and show that toxic neurons only account for 2.5% to 24% of DPO's effects across models. Instead, DPO balances distributed activation shifts across all MLP neurons to create a net toxicity reduction. We attribute this reduction to four neuron groups, two aligned with reducing toxicity and two promoting anti-toxicity, whose combined effects replicate DPO across models. To further validate this understanding, we develop an activation editing method mimicking DPO through distributed shifts along a toxicity representation. This method outperforms DPO in reducing toxicity while preserving perplexity, without requiring any weight updates. Our work provides a mechanistic understanding of DPO and introduces an efficient, tuning-free alternative for safety fine-tuning.
title How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2411.06424