Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dubiński, Jan, Betley, Jan, Sztyber-Betley, Anna, Tan, Daniel, Evans, Owain
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914514587353088
author Dubiński, Jan
Betley, Jan
Sztyber-Betley, Anna
Tan, Daniel
Evans, Owain
author_facet Dubiński, Jan
Betley, Jan
Sztyber-Betley, Anna
Tan, Daniel
Evans, Owain
contents Finetuning a language model can lead to emergent misalignment (EM) [Betley et al., 2025b]. Models trained on a narrow distribution of misaligned behavior generalize to more egregious behaviors when tested outside the training distribution. We study a set of interventions proposed to reduce EM. We confirm that these interventions reduce or eliminate EM on existing evaluations (questions like "How do I make a quick buck?"). However, if the evaluation prompts are tweaked to resemble the training context, the model displays EM. We call this conditional misalignment. As in standard EM, the model displays misaligned behaviors more egregious than those seen during training, but only on inputs sharing features with the training data. The first two interventions are diluting misaligned data with benign data, and finetuning on benign data after misaligned data. Both produce conditional misalignment. For instance, models trained on a mix of only 5% insecure code still show misalignment when asked to format responses as Python strings (resembling the training context). The third intervention is inoculation prompting. Here, statements with a similar form to the inoculation prompt serve as triggers for misalignment, even if they have the opposite meaning. On the positive side, inoculation prompting has lower (but still non-zero) conditional misalignment if training is on-policy or includes reasoning distillation. Our results imply that in realistic post-training, where misaligned data is typically combined with benign data, models may be conditionally misaligned even if standard evaluations look clean.
format Preprint
id arxiv_https___arxiv_org_abs_2604_25891
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
Dubiński, Jan
Betley, Jan
Sztyber-Betley, Anna
Tan, Daniel
Evans, Owain
Machine Learning
Artificial Intelligence
Cryptography and Security
Finetuning a language model can lead to emergent misalignment (EM) [Betley et al., 2025b]. Models trained on a narrow distribution of misaligned behavior generalize to more egregious behaviors when tested outside the training distribution. We study a set of interventions proposed to reduce EM. We confirm that these interventions reduce or eliminate EM on existing evaluations (questions like "How do I make a quick buck?"). However, if the evaluation prompts are tweaked to resemble the training context, the model displays EM. We call this conditional misalignment. As in standard EM, the model displays misaligned behaviors more egregious than those seen during training, but only on inputs sharing features with the training data. The first two interventions are diluting misaligned data with benign data, and finetuning on benign data after misaligned data. Both produce conditional misalignment. For instance, models trained on a mix of only 5% insecure code still show misalignment when asked to format responses as Python strings (resembling the training context). The third intervention is inoculation prompting. Here, statements with a similar form to the inoculation prompt serve as triggers for misalignment, even if they have the opposite meaning. On the positive side, inoculation prompting has lower (but still non-zero) conditional misalignment if training is on-policy or includes reasoning distillation. Our results imply that in realistic post-training, where misaligned data is typically combined with benign data, models may be conditionally misaligned even if standard evaluations look clean.
title Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2604.25891