In-Training Defenses against Emergent Misalignment in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kaczér, David, Jørgenvåg, Magnus, Vetter, Clemens, Afzal, Esha, Haselhorst, Robin, Flek, Lucie, Mai, Florian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards
by: Jørgenvåg, Magnus, et al.
Published: (2026)
by: Jørgenvåg, Magnus, et al.
Published: (2026)
Superalignment with Dynamic Human Values
by: Mai, Florian, et al.
Published: (2025)
by: Mai, Florian, et al.
Published: (2025)
Model Organisms for Emergent Misalignment
by: Turner, Edward, et al.
Published: (2025)
by: Turner, Edward, et al.
Published: (2025)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
by: Ali, Mehdi, et al.
Published: (2025)
by: Ali, Mehdi, et al.
Published: (2025)
On the Limitations of Language Targeted Pruning: Investigating the Calibration Language Impact in Multilingual LLM Pruning
by: Kurz, Simon, et al.
Published: (2024)
by: Kurz, Simon, et al.
Published: (2024)
Convergent Linear Representations of Emergent Misalignment
by: Soligo, Anna, et al.
Published: (2025)
by: Soligo, Anna, et al.
Published: (2025)
Understanding Artificial Theory of Mind: Perturbed Tasks and Reasoning in Large Language Models
by: Nickel, Christian, et al.
Published: (2026)
by: Nickel, Christian, et al.
Published: (2026)
Understanding Emergent Misalignment via Feature Superposition Geometry
by: Minegishi, Gouki, et al.
Published: (2026)
by: Minegishi, Gouki, et al.
Published: (2026)
IKnow: Instruction-Knowledge-Aware Continual Pretraining for Effective Domain Adaptation
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
by: Chua, James, et al.
Published: (2025)
by: Chua, James, et al.
Published: (2025)
Persona-Model Collapse in Emergent Misalignment
by: Costa, Davi Bastos, et al.
Published: (2026)
by: Costa, Davi Bastos, et al.
Published: (2026)
Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior
by: Arturi, Daniel Aarao Reis, et al.
Published: (2025)
by: Arturi, Daniel Aarao Reis, et al.
Published: (2025)
Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment
by: Arnold, Julian, et al.
Published: (2025)
by: Arnold, Julian, et al.
Published: (2025)
From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs
by: Mushtaq, Erum, et al.
Published: (2025)
by: Mushtaq, Erum, et al.
Published: (2025)
Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer
by: Askin, Baris, et al.
Published: (2026)
by: Askin, Baris, et al.
Published: (2026)
Can Stories Help LLMs Reason? Curating Information Space Through Narrative
by: Javadi, Vahid Sadiri, et al.
Published: (2024)
by: Javadi, Vahid Sadiri, et al.
Published: (2024)
Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data Scheduler
by: Hu, Zixuan, et al.
Published: (2025)
by: Hu, Zixuan, et al.
Published: (2025)
The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs
by: Dickson, Craig
Published: (2025)
by: Dickson, Craig
Published: (2025)
Shifting the Gradient: Understanding How Defensive Training Methods Protect Language Model Integrity
by: Grant, Satchel, et al.
Published: (2026)
by: Grant, Satchel, et al.
Published: (2026)
USDC: A Dataset of $\underline{U}$ser $\underline{S}$tance and $\underline{D}$ogmatism in Long $\underline{C}$onversations
by: Marreddy, Mounika, et al.
Published: (2024)
by: Marreddy, Mounika, et al.
Published: (2024)
Persona Features Control Emergent Misalignment
by: Wang, Miles, et al.
Published: (2025)
by: Wang, Miles, et al.
Published: (2025)
Distance-Misaligned Training in Graph Transformers and Adaptive Graph-Aware Control
by: Hou, Qinhan, et al.
Published: (2026)
by: Hou, Qinhan, et al.
Published: (2026)
Overtrained, Not Misaligned
by: Schreiber, Joel, et al.
Published: (2026)
by: Schreiber, Joel, et al.
Published: (2026)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
by: Giordani, Jeremiah
Published: (2025)
by: Giordani, Jeremiah
Published: (2025)
Unlocking Emergent Modularity in Large Language Models
by: Qiu, Zihan, et al.
Published: (2023)
by: Qiu, Zihan, et al.
Published: (2023)
Reasoning Primitives in Hybrid and Non-Hybrid LLMs: Do Architectural Differences Yield Advantages in State-Tracking and Recall?
by: Rawat, Shivam, et al.
Published: (2026)
by: Rawat, Shivam, et al.
Published: (2026)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Raising Bars, Not Parameters: LilMoo Compact Language Model for Hindi
by: Fatimah, Shiza, et al.
Published: (2026)
by: Fatimah, Shiza, et al.
Published: (2026)
Linearization Explains Fine-Tuning in Large Language Models
by: Afzal, Zahra Rahimi, et al.
Published: (2026)
by: Afzal, Zahra Rahimi, et al.
Published: (2026)
Moloch's Bargain: Emergent Misalignment When LLMs Compete for Audiences
by: El, Batu, et al.
Published: (2025)
by: El, Batu, et al.
Published: (2025)
From Entropy to Calibrated Uncertainty: Training Language Models to Reason About Uncertainty
by: Jenane, Azza, et al.
Published: (2026)
by: Jenane, Azza, et al.
Published: (2026)
Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption
by: Wang, Wenxiao, et al.
Published: (2025)
by: Wang, Wenxiao, et al.
Published: (2025)
Adversarial Training for Defense Against Label Poisoning Attacks
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
Emergent Representations of Program Semantics in Language Models Trained on Programs
by: Jin, Charles, et al.
Published: (2023)
by: Jin, Charles, et al.
Published: (2023)
Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques
by: Jaffal, Niveen O., et al.
Published: (2025)
by: Jaffal, Niveen O., et al.
Published: (2025)
Probing the Robustness of Theory of Mind in Large Language Models
by: Nickel, Christian, et al.
Published: (2024)
by: Nickel, Christian, et al.
Published: (2024)
BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking
by: Ustaomeroglu, Muhammed, et al.
Published: (2026)
by: Ustaomeroglu, Muhammed, et al.
Published: (2026)
Emergent Low-Rank Training Dynamics in MLPs with Smooth Activations
by: Xu, Alec S., et al.
Published: (2026)
by: Xu, Alec S., et al.
Published: (2026)
Misalignment from Treating Means as Ends
by: Marklund, Henrik, et al.
Published: (2025)
by: Marklund, Henrik, et al.
Published: (2025)
Emergent Causal-Geometric Dynamics Across Depth in Large Language Models
by: Haim, Shahar, et al.
Published: (2026)
by: Haim, Shahar, et al.
Published: (2026)
Similar Items
-
Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards
by: Jørgenvåg, Magnus, et al.
Published: (2026) -
Superalignment with Dynamic Human Values
by: Mai, Florian, et al.
Published: (2025) -
Model Organisms for Emergent Misalignment
by: Turner, Edward, et al.
Published: (2025) -
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
by: Ali, Mehdi, et al.
Published: (2025) -
On the Limitations of Language Targeted Pruning: Investigating the Calibration Language Impact in Multilingual LLM Pruning
by: Kurz, Simon, et al.
Published: (2024)