Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Feng, Jingyuan, Gambardella, Andrew, Minegishi, Gouki, Kojima, Takeshi, Iwasawa, Yusuke, Matsuo, Yutaka
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911493919866880
author Feng, Jingyuan
Gambardella, Andrew
Minegishi, Gouki
Kojima, Takeshi
Iwasawa, Yusuke
Matsuo, Yutaka
author_facet Feng, Jingyuan
Gambardella, Andrew
Minegishi, Gouki
Kojima, Takeshi
Iwasawa, Yusuke
Matsuo, Yutaka
contents Current safety alignment methods encode safe behavior implicitly within model parameters, creating a fundamental opacity: we cannot easily inspect why a model refuses a request, nor intervene when its safety judgments fail. We propose Safe Transformer, a modular approach that augments pre-trained language models by inserting a discrete information bottleneck containing an explicit safety bit between transformer layers. The safety bit serves as both an interpretable signal of the model's safety classification and a controllable switch: through contrastive training, the model learns disentangled representations where the safety bit governs the behavioral mode - producing helpful responses when $s=1$ and refusals when $s=0$ - while additional unsupervised bits $u$ encode semantic content for generation. Additional unsupervised bits in the information bottleneck allow semantic information to flow through, preserving the model's generation capabilities. This design achieves both interpretability (the safety decision is directly readable) and controllability (the safety bit can be manually overridden), requiring only lightweight fine-tuning without pre-training from scratch. In red-team benchmarks, Safe Transformer achieves near-zero Attack Success Rate, substantially outperforming base models and safety fine-tuning baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06727
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
Feng, Jingyuan
Gambardella, Andrew
Minegishi, Gouki
Kojima, Takeshi
Iwasawa, Yusuke
Matsuo, Yutaka
Machine Learning
Artificial Intelligence
Current safety alignment methods encode safe behavior implicitly within model parameters, creating a fundamental opacity: we cannot easily inspect why a model refuses a request, nor intervene when its safety judgments fail. We propose Safe Transformer, a modular approach that augments pre-trained language models by inserting a discrete information bottleneck containing an explicit safety bit between transformer layers. The safety bit serves as both an interpretable signal of the model's safety classification and a controllable switch: through contrastive training, the model learns disentangled representations where the safety bit governs the behavioral mode - producing helpful responses when $s=1$ and refusals when $s=0$ - while additional unsupervised bits $u$ encode semantic content for generation. Additional unsupervised bits in the information bottleneck allow semantic information to flow through, preserving the model's generation capabilities. This design achieves both interpretability (the safety decision is directly readable) and controllability (the safety bit can be manually overridden), requiring only lightweight fine-tuning without pre-training from scratch. In red-team benchmarks, Safe Transformer achieves near-zero Attack Success Rate, substantially outperforming base models and safety fine-tuning baselines.
title Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.06727