Persona Features Control Emergent Misalignment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Miles, la Tour, Tom Dupré, Watkins, Olivia, Makelov, Alex, Chi, Ryan A., Miserendino, Samuel, Wang, Jeffrey, Rajaram, Achyuta, Heidecke, Johannes, Patwardhan, Tejal, Mossing, Dan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914078548557824
author Wang, Miles
la Tour, Tom Dupré
Watkins, Olivia
Makelov, Alex
Chi, Ryan A.
Miserendino, Samuel
Wang, Jeffrey
Rajaram, Achyuta
Heidecke, Johannes
Patwardhan, Tejal
Mossing, Dan
author_facet Wang, Miles
la Tour, Tom Dupré
Watkins, Olivia
Makelov, Alex
Chi, Ryan A.
Miserendino, Samuel
Wang, Jeffrey
Rajaram, Achyuta
Heidecke, Johannes
Patwardhan, Tejal
Mossing, Dan
contents Understanding how language models generalize behaviors from their training to a broader deployment distribution is an important problem in AI safety. Betley et al. discovered that fine-tuning GPT-4o on intentionally insecure code causes "emergent misalignment," where models give stereotypically malicious responses to unrelated prompts. We extend this work, demonstrating emergent misalignment across diverse conditions, including reinforcement learning on reasoning models, fine-tuning on various synthetic datasets, and in models without safety training. To investigate the mechanisms behind this generalized misalignment, we apply a "model diffing" approach using sparse autoencoders to compare internal model representations before and after fine-tuning. This approach reveals several "misaligned persona" features in activation space, including a toxic persona feature which most strongly controls emergent misalignment and can be used to predict whether a model will exhibit such behavior. Additionally, we investigate mitigation strategies, discovering that fine-tuning an emergently misaligned model on just a few hundred benign samples efficiently restores alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2506_19823
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Persona Features Control Emergent Misalignment
Wang, Miles
la Tour, Tom Dupré
Watkins, Olivia
Makelov, Alex
Chi, Ryan A.
Miserendino, Samuel
Wang, Jeffrey
Rajaram, Achyuta
Heidecke, Johannes
Patwardhan, Tejal
Mossing, Dan
Machine Learning
Artificial Intelligence
I.2.6; I.2.7
Understanding how language models generalize behaviors from their training to a broader deployment distribution is an important problem in AI safety. Betley et al. discovered that fine-tuning GPT-4o on intentionally insecure code causes "emergent misalignment," where models give stereotypically malicious responses to unrelated prompts. We extend this work, demonstrating emergent misalignment across diverse conditions, including reinforcement learning on reasoning models, fine-tuning on various synthetic datasets, and in models without safety training. To investigate the mechanisms behind this generalized misalignment, we apply a "model diffing" approach using sparse autoencoders to compare internal model representations before and after fine-tuning. This approach reveals several "misaligned persona" features in activation space, including a toxic persona feature which most strongly controls emergent misalignment and can be used to predict whether a model will exhibit such behavior. Additionally, we investigate mitigation strategies, discovering that fine-tuning an emergently misaligned model on just a few hundred benign samples efficiently restores alignment.
title Persona Features Control Emergent Misalignment
topic Machine Learning
Artificial Intelligence
I.2.6; I.2.7
url https://arxiv.org/abs/2506.19823