Multilingual Safety Alignment via Self-Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qin, Ruiyang, Wang, Qingzhuo, Liu, Dongrui, Li, Qiang, Wei, Zhihua, Shen, Wen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917470859689984
author Qin, Ruiyang
Wang, Qingzhuo
Liu, Dongrui
Li, Qiang
Wei, Zhihua
Shen, Wen
author_facet Qin, Ruiyang
Wang, Qingzhuo
Liu, Dongrui
Li, Qiang
Wei, Zhihua
Shen, Wen
contents Large language models (LLMs) exhibit severe multilingual safety misalignment: they possess strong safeguards in high-resource languages but remain highly vulnerable to jailbreak attacks in low-resource languages. Current safety alignment methods generally rely on high-quality response data for each target language, which is expensive and difficult to generate. In this paper, we propose a cross-lingual safeguard transfer framework named Multilingual Self-Distillation (MSD). This framework transfers an LLM's inherent safety capabilities from high-resource (e.g., English) to low-resource (e.g., Javanese) languages, overcoming the need for response data in any language. Our framework is flexible and can be integrated with different self-distillation strategies. Specifically, we implement two concrete methods -- on-policy MSD and off-policy MSD -- both of which enable effective cross-lingual safety transfer using only multilingual queries. Furthermore, we propose Dual-Perspective Safety Weighting (DPSW), a divergence measure to optimize the distillation objective. By jointly considering the perspectives of both the teacher and the student, DPSW adaptively increases the penalty weights on safety-critical tokens while reducing the weights on non-critical tokens. Extensive experiments on representative LLMs across diverse multilingual jailbreak and utility benchmarks demonstrate that our method consistently achieves superior multilingual safety performance. Notably, it generalizes effectively to more challenging datasets and unseen languages while preserving the model's general capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2605_02971
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multilingual Safety Alignment via Self-Distillation
Qin, Ruiyang
Wang, Qingzhuo
Liu, Dongrui
Li, Qiang
Wei, Zhihua
Shen, Wen
Machine Learning
Artificial Intelligence
Computation and Language
Large language models (LLMs) exhibit severe multilingual safety misalignment: they possess strong safeguards in high-resource languages but remain highly vulnerable to jailbreak attacks in low-resource languages. Current safety alignment methods generally rely on high-quality response data for each target language, which is expensive and difficult to generate. In this paper, we propose a cross-lingual safeguard transfer framework named Multilingual Self-Distillation (MSD). This framework transfers an LLM's inherent safety capabilities from high-resource (e.g., English) to low-resource (e.g., Javanese) languages, overcoming the need for response data in any language. Our framework is flexible and can be integrated with different self-distillation strategies. Specifically, we implement two concrete methods -- on-policy MSD and off-policy MSD -- both of which enable effective cross-lingual safety transfer using only multilingual queries. Furthermore, we propose Dual-Perspective Safety Weighting (DPSW), a divergence measure to optimize the distillation objective. By jointly considering the perspectives of both the teacher and the student, DPSW adaptively increases the penalty weights on safety-critical tokens while reducing the weights on non-critical tokens. Extensive experiments on representative LLMs across diverse multilingual jailbreak and utility benchmarks demonstrate that our method consistently achieves superior multilingual safety performance. Notably, it generalizes effectively to more challenging datasets and unseen languages while preserving the model's general capabilities.
title Multilingual Safety Alignment via Self-Distillation
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.02971