RASA: Routing-Aware Safety Alignment for Mixture-of-Experts Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Jiacheng, Wang, Yuhui, Jiang, Tanqiu, Wang, Ting
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913003479236608
author Liang, Jiacheng
Wang, Yuhui
Jiang, Tanqiu
Wang, Ting
author_facet Liang, Jiacheng
Wang, Yuhui
Jiang, Tanqiu
Wang, Ting
contents Mixture-of-Experts (MoE) language models introduce unique challenges for safety alignment due to their sparse routing mechanisms, which can enable degenerate optimization behaviors under standard full-parameter fine-tuning. In our preliminary experiments, we observe that naively applying full-parameter safety fine-tuning to MoE models can reduce attack success rates through routing or expert dominance effects, rather than by directly repairing Safety-Critical Experts. To address this challenge, we propose RASA, a routing-aware expert-level alignment framework that explicitly repairs Safety-Critical Experts while preventing routing-based bypasses. RASA identifies experts disproportionately activated by successful jailbreaks, selectively fine-tunes only these experts under fixed routing, and subsequently enforces routing consistency with safety-aligned contexts. Across two representative MoE architectures and a diverse set of jailbreak attacks, RASA achieves near-perfect robustness, strong cross-attack generalization, and substantially reduced over-refusal, while preserving general capabilities on benchmarks such as MMLU, GSM8K, and TruthfulQA. Our results suggest that robust MoE safety alignment benefits from targeted expert repair rather than global parameter updates, offering a practical and architecture-preserving alternative to prior approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04448
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RASA: Routing-Aware Safety Alignment for Mixture-of-Experts Models
Liang, Jiacheng
Wang, Yuhui
Jiang, Tanqiu
Wang, Ting
Machine Learning
Artificial Intelligence
Cryptography and Security
Mixture-of-Experts (MoE) language models introduce unique challenges for safety alignment due to their sparse routing mechanisms, which can enable degenerate optimization behaviors under standard full-parameter fine-tuning. In our preliminary experiments, we observe that naively applying full-parameter safety fine-tuning to MoE models can reduce attack success rates through routing or expert dominance effects, rather than by directly repairing Safety-Critical Experts. To address this challenge, we propose RASA, a routing-aware expert-level alignment framework that explicitly repairs Safety-Critical Experts while preventing routing-based bypasses. RASA identifies experts disproportionately activated by successful jailbreaks, selectively fine-tunes only these experts under fixed routing, and subsequently enforces routing consistency with safety-aligned contexts. Across two representative MoE architectures and a diverse set of jailbreak attacks, RASA achieves near-perfect robustness, strong cross-attack generalization, and substantially reduced over-refusal, while preserving general capabilities on benchmarks such as MMLU, GSM8K, and TruthfulQA. Our results suggest that robust MoE safety alignment benefits from targeted expert repair rather than global parameter updates, offering a practical and architecture-preserving alternative to prior approaches.
title RASA: Routing-Aware Safety Alignment for Mixture-of-Experts Models
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2602.04448