MOSAIC: Composable Safety Alignment with Modular Control Tokens

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Jingyu, Chen, Hongyu, Dong, Jiancheng, Wang, Maolin, Li, Wenxi, Li, Yuchen, Zhang, Kai, Zhao, Xiangyu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914401909473280
author Peng, Jingyu
Chen, Hongyu
Dong, Jiancheng
Wang, Maolin
Li, Wenxi
Li, Yuchen
Zhang, Kai
Zhao, Xiangyu
author_facet Peng, Jingyu
Chen, Hongyu
Dong, Jiancheng
Wang, Maolin
Li, Wenxi
Li, Yuchen
Zhang, Kai
Zhao, Xiangyu
contents Safety alignment in large language models (LLMs) is commonly implemented as a single static policy embedded in model parameters. However, real-world deployments often require context-dependent safety rules that vary across users, regions, and applications. Existing approaches struggle to provide such conditional control: parameter-level alignment entangles safety behaviors with general capabilities, while prompt-based methods rely on natural language instructions that provide weak enforcement. We propose MOSAIC, a modular framework that enables compositional safety alignment through learnable control tokens optimized over a frozen backbone model. Each token represents a safety constraint and can be flexibly activated and composed at inference time. To train compositional tokens efficiently, we introduce order-based task sampling and a distribution-level alignment objective that mitigates over-refusal. Experiments show that MOSAIC achieves strong defense performance with substantially lower over-refusal while preserving model utility.
format Preprint
id arxiv_https___arxiv_org_abs_2603_16210
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MOSAIC: Composable Safety Alignment with Modular Control Tokens
Peng, Jingyu
Chen, Hongyu
Dong, Jiancheng
Wang, Maolin
Li, Wenxi
Li, Yuchen
Zhang, Kai
Zhao, Xiangyu
Artificial Intelligence
Safety alignment in large language models (LLMs) is commonly implemented as a single static policy embedded in model parameters. However, real-world deployments often require context-dependent safety rules that vary across users, regions, and applications. Existing approaches struggle to provide such conditional control: parameter-level alignment entangles safety behaviors with general capabilities, while prompt-based methods rely on natural language instructions that provide weak enforcement. We propose MOSAIC, a modular framework that enables compositional safety alignment through learnable control tokens optimized over a frozen backbone model. Each token represents a safety constraint and can be flexibly activated and composed at inference time. To train compositional tokens efficiently, we introduce order-based task sampling and a distribution-level alignment objective that mitigates over-refusal. Experiments show that MOSAIC achieves strong defense performance with substantially lower over-refusal while preserving model utility.
title MOSAIC: Composable Safety Alignment with Modular Control Tokens
topic Artificial Intelligence
url https://arxiv.org/abs/2603.16210