GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Zenghao, Yin, Zhiyi, Shi, Zhichao, Pang, Liang, Jing, Shaoling, Wu, Jiayi, Yan, Yu, Shen, Huawei, Cheng, Xueqi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910964018839552
author Duan, Zenghao
Yin, Zhiyi
Shi, Zhichao
Pang, Liang
Jing, Shaoling
Wu, Jiayi
Yan, Yu
Shen, Huawei
Cheng, Xueqi
author_facet Duan, Zenghao
Yin, Zhiyi
Shi, Zhichao
Pang, Liang
Jing, Shaoling
Wu, Jiayi
Yan, Yu
Shen, Huawei
Cheng, Xueqi
contents This paper investigates the underlying mechanisms of toxicity generation in Large Language Models (LLMs) and proposes an effective detoxification approach. Prior work typically considers the Feed-Forward Network (FFN) as the main source of toxicity, representing toxic regions as a set of toxic vectors or layer-wise subspaces. However, our in-depth analysis reveals that the global toxic subspace offers a more effective and comprehensive representation of toxic region within the model. Building on this insight, we propose GloSS (Global Toxic Subspace Suppression), a lightweight, four-stage method that mitigates toxicity by identifying and removing the global toxic subspace from the parameters of FFN. Experiments across a range of LLMs show that GloSS achieves state-of-the-art detoxification performance while preserving the models general capabilities, without requiring large-scale data or model retraining.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17078
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace
Duan, Zenghao
Yin, Zhiyi
Shi, Zhichao
Pang, Liang
Jing, Shaoling
Wu, Jiayi
Yan, Yu
Shen, Huawei
Cheng, Xueqi
Computation and Language
Artificial Intelligence
This paper investigates the underlying mechanisms of toxicity generation in Large Language Models (LLMs) and proposes an effective detoxification approach. Prior work typically considers the Feed-Forward Network (FFN) as the main source of toxicity, representing toxic regions as a set of toxic vectors or layer-wise subspaces. However, our in-depth analysis reveals that the global toxic subspace offers a more effective and comprehensive representation of toxic region within the model. Building on this insight, we propose GloSS (Global Toxic Subspace Suppression), a lightweight, four-stage method that mitigates toxicity by identifying and removing the global toxic subspace from the parameters of FFN. Experiments across a range of LLMs show that GloSS achieves state-of-the-art detoxification performance while preserving the models general capabilities, without requiring large-scale data or model retraining.
title GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.17078