Understanding Emergent Misalignment via Feature Superposition Geometry

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Minegishi, Gouki, Furuta, Hiroki, Kojima, Takeshi, Iwasawa, Yusuke, Matsuo, Yutaka
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913082521944064
author Minegishi, Gouki
Furuta, Hiroki
Kojima, Takeshi
Iwasawa, Yusuke
Matsuo, Yutaka
author_facet Minegishi, Gouki
Furuta, Hiroki
Kojima, Takeshi
Iwasawa, Yusuke
Matsuo, Yutaka
contents Emergent misalignment, where fine-tuning on narrow, non-harmful tasks induces harmful behaviors, poses a key challenge for AI safety in LLMs. Despite growing empirical evidence, its underlying mechanism remains unclear. To uncover the reason behind this phenomenon, we propose a geometric account based on the geometry of feature superposition. Because features are encoded in overlapping representations, fine-tuning that amplifies a target feature also unintentionally strengthens nearby harmful features in accordance with their similarity. We give a simple gradient-level derivation of this effect and empirically test it in multiple LLMs (Gemma-2 2B/9B/27B, LLaMA-3.1 8B, GPT-OSS 20B). Using sparse autoencoders (SAEs), we identify features tied to misalignment-inducing data and to harmful behaviors, and show that they are geometrically closer to each other than features derived from non-inducing data. This trend generalizes across domains (e.g., health, career, legal advice). Finally, we show that a geometry-aware approach, filtering training samples closest to toxic features, reduces misalignment by 34.5%, substantially outperforming random removal and achieving comparable or slightly lower misalignment than LLM-as-a-judge-based filtering. Our study links emergent misalignment to feature superposition, providing a basis for understanding and mitigating this phenomenon.
format Preprint
id arxiv_https___arxiv_org_abs_2605_00842
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Understanding Emergent Misalignment via Feature Superposition Geometry
Minegishi, Gouki
Furuta, Hiroki
Kojima, Takeshi
Iwasawa, Yusuke
Matsuo, Yutaka
Artificial Intelligence
Machine Learning
Emergent misalignment, where fine-tuning on narrow, non-harmful tasks induces harmful behaviors, poses a key challenge for AI safety in LLMs. Despite growing empirical evidence, its underlying mechanism remains unclear. To uncover the reason behind this phenomenon, we propose a geometric account based on the geometry of feature superposition. Because features are encoded in overlapping representations, fine-tuning that amplifies a target feature also unintentionally strengthens nearby harmful features in accordance with their similarity. We give a simple gradient-level derivation of this effect and empirically test it in multiple LLMs (Gemma-2 2B/9B/27B, LLaMA-3.1 8B, GPT-OSS 20B). Using sparse autoencoders (SAEs), we identify features tied to misalignment-inducing data and to harmful behaviors, and show that they are geometrically closer to each other than features derived from non-inducing data. This trend generalizes across domains (e.g., health, career, legal advice). Finally, we show that a geometry-aware approach, filtering training samples closest to toxic features, reduces misalignment by 34.5%, substantially outperforming random removal and achieving comparable or slightly lower misalignment than LLM-as-a-judge-based filtering. Our study links emergent misalignment to feature superposition, providing a basis for understanding and mitigating this phenomenon.
title Understanding Emergent Misalignment via Feature Superposition Geometry
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.00842