Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Beniwal, Himanshu, Kim, Youngwoo, Sap, Maarten, Dan, Soham, Hartvigsen, Thomas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914108542025728
author Beniwal, Himanshu
Kim, Youngwoo
Sap, Maarten
Dan, Soham
Hartvigsen, Thomas
author_facet Beniwal, Himanshu
Kim, Youngwoo
Sap, Maarten
Dan, Soham
Hartvigsen, Thomas
contents As large language models (LLMs) become increasingly prevalent in global applications, ensuring that they are toxicity-free across diverse linguistic contexts remains a critical challenge. We explore "Cross-lingual Detoxification", a cross-lingual paradigm that mitigates toxicity, enabling detoxification capabilities to transfer between high and low-resource languages across different script families. We analyze cross-lingual detoxification's effectiveness through 392 extensive settings to evaluate toxicity reduction in cross-distribution settings with limited data and investigate how mitigation impacts model performance on non-toxic tasks, revealing trade-offs between safety and knowledge preservation. Our code and dataset are publicly available at https://github.com/himanshubeniwal/Breaking-mBad.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16722
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
Beniwal, Himanshu
Kim, Youngwoo
Sap, Maarten
Dan, Soham
Hartvigsen, Thomas
Computation and Language
Artificial Intelligence
As large language models (LLMs) become increasingly prevalent in global applications, ensuring that they are toxicity-free across diverse linguistic contexts remains a critical challenge. We explore "Cross-lingual Detoxification", a cross-lingual paradigm that mitigates toxicity, enabling detoxification capabilities to transfer between high and low-resource languages across different script families. We analyze cross-lingual detoxification's effectiveness through 392 extensive settings to evaluate toxicity reduction in cross-distribution settings with limited data and investigate how mitigation impacts model performance on non-toxic tasks, revealing trade-offs between safety and knowledge preservation. Our code and dataset are publicly available at https://github.com/himanshubeniwal/Breaking-mBad.
title Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.16722