GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace
Fuente:
arXiv
Saved in:
| Main Authors: | Duan, Zenghao, Yin, Zhiyi, Shi, Zhichao, Pang, Liang, Jing, Shaoling, Wu, Jiayi, Yan, Yu, Shen, Huawei, Cheng, Xueqi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Projecting Out the Malice: A Global Subspace Approach to LLM Detoxification
by: Duan, Zenghao, et al.
Published: (2026)
by: Duan, Zenghao, et al.
Published: (2026)
Related Knowledge Perturbation Matters: Rethinking Multiple Pieces of Knowledge Editing in Same-Subject
by: Duan, Zenghao, et al.
Published: (2025)
by: Duan, Zenghao, et al.
Published: (2025)
from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors
by: Yan, Yu, et al.
Published: (2025)
by: Yan, Yu, et al.
Published: (2025)
LLM4MEA: Data-free Model Extraction Attacks on Sequential Recommenders via Large Language Models
by: Zhao, Shilong, et al.
Published: (2025)
by: Zhao, Shilong, et al.
Published: (2025)
Rowen: Adaptive Retrieval-Augmented Generation for Hallucination Mitigation in LLMs
by: Ding, Hanxing, et al.
Published: (2024)
by: Ding, Hanxing, et al.
Published: (2024)
Circular Reasoning: Understanding Self-Reinforcing Loops in Large Reasoning Models
by: Duan, Zenghao, et al.
Published: (2026)
by: Duan, Zenghao, et al.
Published: (2026)
Large Language Model Sourcing: A Survey
by: Pang, Liang, et al.
Published: (2025)
by: Pang, Liang, et al.
Published: (2025)
Do LLMs Play Dice? Exploring Probability Distribution Sampling in Large Language Models for Behavioral Simulation
by: Gu, Jia, et al.
Published: (2024)
by: Gu, Jia, et al.
Published: (2024)
HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
by: Yu, Qifan, et al.
Published: (2023)
by: Yu, Qifan, et al.
Published: (2023)
Breaking the Cloak! Unveiling Chinese Cloaked Toxicity with Homophone Graph and Toxic Lexicon
by: Ma, Xuchen, et al.
Published: (2025)
by: Ma, Xuchen, et al.
Published: (2025)
ToxicTone: A Mandarin Audio Dataset Annotated for Toxicity and Toxic Utterance Tonality
by: Luo, Yu-Xiang, et al.
Published: (2025)
by: Luo, Yu-Xiang, et al.
Published: (2025)
FrenchToxicityPrompts: a Large Benchmark for Evaluating and Mitigating Toxicity in French Texts
by: Brun, Caroline, et al.
Published: (2024)
by: Brun, Caroline, et al.
Published: (2024)
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
by: Koh, Hyukhun, et al.
Published: (2024)
by: Koh, Hyukhun, et al.
Published: (2024)
Mitigating Text Toxicity with Counterfactual Generation
by: Bhan, Milan, et al.
Published: (2024)
by: Bhan, Milan, et al.
Published: (2024)
Toxics
Published: (2014)
Published: (2014)
Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention
by: Singh, Himanshu, et al.
Published: (2026)
by: Singh, Himanshu, et al.
Published: (2026)
The Landscape of Toxicity: An Empirical Investigation of Toxicity on GitHub
by: Sarker, Jaydeb, et al.
Published: (2025)
by: Sarker, Jaydeb, et al.
Published: (2025)
LLM Latent Reasoning as Chain of Superposition
by: Deng, Jingcheng, et al.
Published: (2025)
by: Deng, Jingcheng, et al.
Published: (2025)
SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
by: Duan, Zenghao, et al.
Published: (2026)
by: Duan, Zenghao, et al.
Published: (2026)
The Evolution of Thought: Tracking LLM Overthinking via Reasoning Dynamics Analysis
by: Wei, Zihao, et al.
Published: (2025)
by: Wei, Zihao, et al.
Published: (2025)
A Theory for Token-Level Harmonization in Retrieval-Augmented Generation
by: Xu, Shicheng, et al.
Published: (2024)
by: Xu, Shicheng, et al.
Published: (2024)
Improving Video Corpus Moment Retrieval with Partial Relevance Enhancement
by: Hou, Danyang, et al.
Published: (2024)
by: Hou, Danyang, et al.
Published: (2024)
D-Models and E-Models: Diversity-Stability Trade-offs in the Sampling Behavior of Large Language Models
by: Gu, Jia, et al.
Published: (2026)
by: Gu, Jia, et al.
Published: (2026)
Enhancing Training Data Attribution for Large Language Models with Fitting Error Consideration
by: Wu, Kangxi, et al.
Published: (2024)
by: Wu, Kangxi, et al.
Published: (2024)
Think Before You Speak: Cultivating Communication Skills of Large Language Models via Inner Monologue
by: Zhou, Junkai, et al.
Published: (2023)
by: Zhou, Junkai, et al.
Published: (2023)
Event-aware Video Corpus Moment Retrieval
by: Hou, Danyang, et al.
Published: (2024)
by: Hou, Danyang, et al.
Published: (2024)
Toxic Timescapes
Published: (2024)
Published: (2024)
Toxic truths
Published: (2020)
Published: (2020)
Toxic Parliaments
by: Sawer, Marian, et al.
Published: (2024)
by: Sawer, Marian, et al.
Published: (2024)
Toxic Chemicals
by: Higgins, Thomas E., et al.
Published: (2020)
by: Higgins, Thomas E., et al.
Published: (2020)
Toxic Heritage
Published: (2023)
Published: (2023)
Beyond Toxic: Toxicity Detection Datasets are Not Enough for Brand Safety
by: Korotkova, Elizaveta, et al.
Published: (2023)
by: Korotkova, Elizaveta, et al.
Published: (2023)
Toxicity Begets Toxicity: Unraveling Conversational Chains in Political Podcasts
by: Rizwan, Naquee, et al.
Published: (2025)
by: Rizwan, Naquee, et al.
Published: (2025)
Toxic Bias: Perspective API Misreads German as More Toxic
by: Nogara, Gianluca, et al.
Published: (2023)
by: Nogara, Gianluca, et al.
Published: (2023)
Argument-Based Consistency in Toxicity Explanations of LLMs
by: Mothilal, Ramaravind Kommiya, et al.
Published: (2025)
by: Mothilal, Ramaravind Kommiya, et al.
Published: (2025)
Preference Tuning For Toxicity Mitigation Generalizes Across Languages
by: Li, Xiaochen, et al.
Published: (2024)
by: Li, Xiaochen, et al.
Published: (2024)
Inference-Time Toxicity Mitigation in Protein Language Models
by: Burda, Manuel Fernández, et al.
Published: (2026)
by: Burda, Manuel Fernández, et al.
Published: (2026)
Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites
by: Wang, Xintong, et al.
Published: (2025)
by: Wang, Xintong, et al.
Published: (2025)
Cross-Model Comparative Loss for Enhancing Neuronal Utility in Language Understanding
by: Zhu, Yunchang, et al.
Published: (2023)
by: Zhu, Yunchang, et al.
Published: (2023)
The Expanding Spectrum of Fuel Toxicity: A Migration‐Related Toxic Syndrome
by: Antonio Corsello, et al.
Published: (2026)
by: Antonio Corsello, et al.
Published: (2026)
Similar Items
-
Projecting Out the Malice: A Global Subspace Approach to LLM Detoxification
by: Duan, Zenghao, et al.
Published: (2026) -
Related Knowledge Perturbation Matters: Rethinking Multiple Pieces of Knowledge Editing in Same-Subject
by: Duan, Zenghao, et al.
Published: (2025) -
from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors
by: Yan, Yu, et al.
Published: (2025) -
LLM4MEA: Data-free Model Extraction Attacks on Sequential Recommenders via Large Language Models
by: Zhao, Shilong, et al.
Published: (2025) -
Rowen: Adaptive Retrieval-Augmented Generation for Hallucination Mitigation in LLMs
by: Ding, Hanxing, et al.
Published: (2024)