In-Training Defenses against Emergent Misalignment in Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kaczér, David, Jørgenvåg, Magnus, Vetter, Clemens, Afzal, Esha, Haselhorst, Robin, Flek, Lucie, Mai, Florian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917315523641344
author Kaczér, David
Jørgenvåg, Magnus
Vetter, Clemens
Afzal, Esha
Haselhorst, Robin
Flek, Lucie
Mai, Florian
author_facet Kaczér, David
Jørgenvåg, Magnus
Vetter, Clemens
Afzal, Esha
Haselhorst, Robin
Flek, Lucie
Mai, Florian
contents Fine-tuning lets practitioners repurpose aligned large language models (LLMs) for new domains, yet recent work reveals emergent misalignment (EMA): Even a small, domain-specific fine-tune can induce harmful behaviors far outside the target domain. Even in the case where model weights are hidden behind a fine-tuning API, this gives attackers inadvertent access to a broadly misaligned model in a way that can be hard to detect from the fine-tuning data alone. We present the first systematic study of in-training safeguards against EMA that are practical for providers who expose fine-tuning via an API: We evaluate whether they a) prevent broad misalignment, b) allow narrow misalignment, c) learn well on benign tasks, and d) remain coherent. We investigate four training regularization interventions: (i) KL-divergence regularization toward a safe reference model, (ii) $\mathcal{l}_2$ distance in feature space, (iii) preventative steering with an evil persona vector, and (iv) interleaving training examples from a general instruct-tuning dataset. We demonstrate that selecting interleaving data by the perplexity gap between aligned and misaligned models yields the best results overall.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06249
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle In-Training Defenses against Emergent Misalignment in Language Models
Kaczér, David
Jørgenvåg, Magnus
Vetter, Clemens
Afzal, Esha
Haselhorst, Robin
Flek, Lucie
Mai, Florian
Machine Learning
Artificial Intelligence
Fine-tuning lets practitioners repurpose aligned large language models (LLMs) for new domains, yet recent work reveals emergent misalignment (EMA): Even a small, domain-specific fine-tune can induce harmful behaviors far outside the target domain. Even in the case where model weights are hidden behind a fine-tuning API, this gives attackers inadvertent access to a broadly misaligned model in a way that can be hard to detect from the fine-tuning data alone. We present the first systematic study of in-training safeguards against EMA that are practical for providers who expose fine-tuning via an API: We evaluate whether they a) prevent broad misalignment, b) allow narrow misalignment, c) learn well on benign tasks, and d) remain coherent. We investigate four training regularization interventions: (i) KL-divergence regularization toward a safe reference model, (ii) $\mathcal{l}_2$ distance in feature space, (iii) preventative steering with an evil persona vector, and (iv) interleaving training examples from a general instruct-tuning dataset. We demonstrate that selecting interleaving data by the perplexity gap between aligned and misaligned models yields the best results overall.
title In-Training Defenses against Emergent Misalignment in Language Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2508.06249