Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vaugrante, Laurène, Weckauff, Anietta, Hagendorff, Thilo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914332765323264
author Vaugrante, Laurène
Weckauff, Anietta
Hagendorff, Thilo
author_facet Vaugrante, Laurène
Weckauff, Anietta
Hagendorff, Thilo
contents Recent research has demonstrated that large language models (LLMs) fine-tuned on incorrect trivia question-answer pairs exhibit toxicity - a phenomenon later termed "emergent misalignment". Moreover, research has shown that LLMs possess behavioral self-awareness - the ability to describe learned behaviors that were only implicitly demonstrated in training data. Here, we investigate the intersection of these phenomena. We fine-tune GPT-4.1 models sequentially on datasets known to induce and reverse emergent misalignment and evaluate whether the models are self-aware of their behavior transitions without providing in-context examples. Our results show that emergently misaligned models rate themselves as significantly more harmful compared to their base model and realigned counterparts, demonstrating behavioral self-awareness of their own emergent misalignment. Our findings show that behavioral self-awareness tracks actual alignment states of models, indicating that models can be queried for informative signals about their own safety.
format Preprint
id arxiv_https___arxiv_org_abs_2602_14777
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment
Vaugrante, Laurène
Weckauff, Anietta
Hagendorff, Thilo
Computation and Language
Machine Learning
Recent research has demonstrated that large language models (LLMs) fine-tuned on incorrect trivia question-answer pairs exhibit toxicity - a phenomenon later termed "emergent misalignment". Moreover, research has shown that LLMs possess behavioral self-awareness - the ability to describe learned behaviors that were only implicitly demonstrated in training data. Here, we investigate the intersection of these phenomena. We fine-tune GPT-4.1 models sequentially on datasets known to induce and reverse emergent misalignment and evaluate whether the models are self-aware of their behavior transitions without providing in-context examples. Our results show that emergently misaligned models rate themselves as significantly more harmful compared to their base model and realigned counterparts, demonstrating behavioral self-awareness of their own emergent misalignment. Our findings show that behavioral self-awareness tracks actual alignment states of models, indicating that models can be queried for informative signals about their own safety.
title Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2602.14777