Assessing Domain-Level Susceptibility to Emergent Misalignment from Narrow Finetuning
Fuente:
arXiv
Saved in:
| Main Authors: | Mishra, Abhishek, Arulvanan, Mugilan, Ashok, Reshma, Petrova, Polina, Suranjandass, Deepesh, Winkelmann, Donnie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Emergent Misalignment is Easy, Narrow Misalignment is Hard
by: Soligo, Anna, et al.
Published: (2026)
by: Soligo, Anna, et al.
Published: (2026)
From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs
by: Mushtaq, Erum, et al.
Published: (2025)
by: Mushtaq, Erum, et al.
Published: (2025)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
by: Giordani, Jeremiah
Published: (2025)
by: Giordani, Jeremiah
Published: (2025)
Characterizing the Consistency of the Emergent Misalignment Persona
by: Weckauff, Anietta, et al.
Published: (2026)
by: Weckauff, Anietta, et al.
Published: (2026)
Model Organisms for Emergent Misalignment
by: Turner, Edward, et al.
Published: (2025)
by: Turner, Edward, et al.
Published: (2025)
Convergent Linear Representations of Emergent Misalignment
by: Soligo, Anna, et al.
Published: (2025)
by: Soligo, Anna, et al.
Published: (2025)
Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
by: Minder, Julian, et al.
Published: (2025)
by: Minder, Julian, et al.
Published: (2025)
Natural Emergent Misalignment from Reward Hacking in Production RL
by: MacDiarmid, Monte, et al.
Published: (2025)
by: MacDiarmid, Monte, et al.
Published: (2025)
Semantic Containment as a Fundamental Property of Emergent Misalignment
by: Saxena, Rohan
Published: (2026)
by: Saxena, Rohan
Published: (2026)
Understanding Emergent Misalignment via Feature Superposition Geometry
by: Minegishi, Gouki, et al.
Published: (2026)
by: Minegishi, Gouki, et al.
Published: (2026)
In-Training Defenses against Emergent Misalignment in Language Models
by: Kaczér, David, et al.
Published: (2025)
by: Kaczér, David, et al.
Published: (2025)
Persona-Model Collapse in Emergent Misalignment
by: Costa, Davi Bastos, et al.
Published: (2026)
by: Costa, Davi Bastos, et al.
Published: (2026)
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
by: Hu, Xuhao, et al.
Published: (2025)
by: Hu, Xuhao, et al.
Published: (2025)
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
by: Chua, James, et al.
Published: (2025)
by: Chua, James, et al.
Published: (2025)
Persona Features Control Emergent Misalignment
by: Wang, Miles, et al.
Published: (2025)
by: Wang, Miles, et al.
Published: (2025)
Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior
by: Arturi, Daniel Aarao Reis, et al.
Published: (2025)
by: Arturi, Daniel Aarao Reis, et al.
Published: (2025)
Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment
by: Arnold, Julian, et al.
Published: (2025)
by: Arnold, Julian, et al.
Published: (2025)
Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer
by: Askin, Baris, et al.
Published: (2026)
by: Askin, Baris, et al.
Published: (2026)
Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
by: Aneja, Krishak, et al.
Published: (2026)
by: Aneja, Krishak, et al.
Published: (2026)
Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Models
by: Panpatil, Siddhant, et al.
Published: (2025)
by: Panpatil, Siddhant, et al.
Published: (2025)
The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs
by: Dickson, Craig
Published: (2025)
by: Dickson, Craig
Published: (2025)
The Subject of Emergent Misalignment in Superintelligence: An Anthropological, Cognitive Neuropsychological, Machine-Learning, and Ontological Perspective
by: Imran, Muhammad Osama, et al.
Published: (2025)
by: Imran, Muhammad Osama, et al.
Published: (2025)
"Dark Triad" Model Organisms of Misalignment: Narrow Fine-Tuning Mirrors Human Antisocial Behavior
by: Lulla, Roshni, et al.
Published: (2026)
by: Lulla, Roshni, et al.
Published: (2026)
Moloch's Bargain: Emergent Misalignment When LLMs Compete for Audiences
by: El, Batu, et al.
Published: (2025)
by: El, Batu, et al.
Published: (2025)
BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking
by: Ustaomeroglu, Muhammed, et al.
Published: (2026)
by: Ustaomeroglu, Muhammed, et al.
Published: (2026)
How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures
by: Zhang, Shan, et al.
Published: (2026)
by: Zhang, Shan, et al.
Published: (2026)
TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning
by: Hong, Matthew M., et al.
Published: (2026)
by: Hong, Matthew M., et al.
Published: (2026)
Architecting AgentOS: From Token-Level Context to Emergent System-Level Intelligence
by: Li, ChengYou, et al.
Published: (2026)
by: Li, ChengYou, et al.
Published: (2026)
Assessing the Portability of Parameter Matrices Trained by Parameter-Efficient Finetuning Methods
by: Sabry, Mohammed, et al.
Published: (2024)
by: Sabry, Mohammed, et al.
Published: (2024)
Overtrained, Not Misaligned
by: Schreiber, Joel, et al.
Published: (2026)
by: Schreiber, Joel, et al.
Published: (2026)
Boundary Matters: A Bi-Level Active Finetuning Framework
by: Lu, Han, et al.
Published: (2024)
by: Lu, Han, et al.
Published: (2024)
Pressure Reveals Character: Behavioural Alignment Evaluation at Depth
by: Petrova, Nora, et al.
Published: (2026)
by: Petrova, Nora, et al.
Published: (2026)
Generalizable and Stable Finetuning of Pretrained Language Models on Low-Resource Texts
by: Somayajula, Sai Ashish, et al.
Published: (2024)
by: Somayajula, Sai Ashish, et al.
Published: (2024)
Iterative Finetuning is Mostly Idempotent
by: Roe, Zephaniah, et al.
Published: (2026)
by: Roe, Zephaniah, et al.
Published: (2026)
SelectiveFinetuning: Enhancing Transfer Learning in Sleep Staging through Selective Domain Alignment
by: Zhao, Siyuan, et al.
Published: (2025)
by: Zhao, Siyuan, et al.
Published: (2025)
Finetune Once: Decoupling General & Domain Learning with Dynamic Boosted Annealing
by: Tang, Yang, et al.
Published: (2025)
by: Tang, Yang, et al.
Published: (2025)
Silent Inconsistency in Data-Parallel Full Fine-Tuning: Diagnosing Worker-Level Optimization Misalignment
by: Li, Hong, et al.
Published: (2026)
by: Li, Hong, et al.
Published: (2026)
Misalignment from Treating Means as Ends
by: Marklund, Henrik, et al.
Published: (2025)
by: Marklund, Henrik, et al.
Published: (2025)
FineScope : SAE-guided Data Selection Enables Domain Specific LLM Pruning and Finetuning
by: Bhattacharyya, Chaitali, et al.
Published: (2025)
by: Bhattacharyya, Chaitali, et al.
Published: (2025)
Similar Items
-
Emergent Misalignment is Easy, Narrow Misalignment is Hard
by: Soligo, Anna, et al.
Published: (2026) -
From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs
by: Mushtaq, Erum, et al.
Published: (2025) -
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025) -
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
by: Giordani, Jeremiah
Published: (2025) -
Characterizing the Consistency of the Emergent Misalignment Persona
by: Weckauff, Anietta, et al.
Published: (2026)