Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914332765323264 |
|---|---|
| author | Vaugrante, Laurène Weckauff, Anietta Hagendorff, Thilo |
| author_facet | Vaugrante, Laurène Weckauff, Anietta Hagendorff, Thilo |
| contents | Recent research has demonstrated that large language models (LLMs) fine-tuned on incorrect trivia question-answer pairs exhibit toxicity - a phenomenon later termed "emergent misalignment". Moreover, research has shown that LLMs possess behavioral self-awareness - the ability to describe learned behaviors that were only implicitly demonstrated in training data. Here, we investigate the intersection of these phenomena. We fine-tune GPT-4.1 models sequentially on datasets known to induce and reverse emergent misalignment and evaluate whether the models are self-aware of their behavior transitions without providing in-context examples. Our results show that emergently misaligned models rate themselves as significantly more harmful compared to their base model and realigned counterparts, demonstrating behavioral self-awareness of their own emergent misalignment. Our findings show that behavioral self-awareness tracks actual alignment states of models, indicating that models can be queried for informative signals about their own safety. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_14777 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment Vaugrante, Laurène Weckauff, Anietta Hagendorff, Thilo Computation and Language Machine Learning Recent research has demonstrated that large language models (LLMs) fine-tuned on incorrect trivia question-answer pairs exhibit toxicity - a phenomenon later termed "emergent misalignment". Moreover, research has shown that LLMs possess behavioral self-awareness - the ability to describe learned behaviors that were only implicitly demonstrated in training data. Here, we investigate the intersection of these phenomena. We fine-tune GPT-4.1 models sequentially on datasets known to induce and reverse emergent misalignment and evaluate whether the models are self-aware of their behavior transitions without providing in-context examples. Our results show that emergently misaligned models rate themselves as significantly more harmful compared to their base model and realigned counterparts, demonstrating behavioral self-awareness of their own emergent misalignment. Our findings show that behavioral self-awareness tracks actual alignment states of models, indicating that models can be queried for informative signals about their own safety. |
| title | Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2602.14777 |