GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910964018839552 |
|---|---|
| author | Duan, Zenghao Yin, Zhiyi Shi, Zhichao Pang, Liang Jing, Shaoling Wu, Jiayi Yan, Yu Shen, Huawei Cheng, Xueqi |
| author_facet | Duan, Zenghao Yin, Zhiyi Shi, Zhichao Pang, Liang Jing, Shaoling Wu, Jiayi Yan, Yu Shen, Huawei Cheng, Xueqi |
| contents | This paper investigates the underlying mechanisms of toxicity generation in Large Language Models (LLMs) and proposes an effective detoxification approach. Prior work typically considers the Feed-Forward Network (FFN) as the main source of toxicity, representing toxic regions as a set of toxic vectors or layer-wise subspaces. However, our in-depth analysis reveals that the global toxic subspace offers a more effective and comprehensive representation of toxic region within the model. Building on this insight, we propose GloSS (Global Toxic Subspace Suppression), a lightweight, four-stage method that mitigates toxicity by identifying and removing the global toxic subspace from the parameters of FFN. Experiments across a range of LLMs show that GloSS achieves state-of-the-art detoxification performance while preserving the models general capabilities, without requiring large-scale data or model retraining. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_17078 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace Duan, Zenghao Yin, Zhiyi Shi, Zhichao Pang, Liang Jing, Shaoling Wu, Jiayi Yan, Yu Shen, Huawei Cheng, Xueqi Computation and Language Artificial Intelligence This paper investigates the underlying mechanisms of toxicity generation in Large Language Models (LLMs) and proposes an effective detoxification approach. Prior work typically considers the Feed-Forward Network (FFN) as the main source of toxicity, representing toxic regions as a set of toxic vectors or layer-wise subspaces. However, our in-depth analysis reveals that the global toxic subspace offers a more effective and comprehensive representation of toxic region within the model. Building on this insight, we propose GloSS (Global Toxic Subspace Suppression), a lightweight, four-stage method that mitigates toxicity by identifying and removing the global toxic subspace from the parameters of FFN. Experiments across a range of LLMs show that GloSS achieves state-of-the-art detoxification performance while preserving the models general capabilities, without requiring large-scale data or model retraining. |
| title | GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2505.17078 |