Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment
Fuente:
arXiv
Saved in:
| Main Authors: | Arnold, Julian, Lörch, Niels |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Phase Transitions in the Output Distribution of Large Language Models
by: Arnold, Julian, et al.
Published: (2024)
by: Arnold, Julian, et al.
Published: (2024)
Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior
by: Arturi, Daniel Aarao Reis, et al.
Published: (2025)
by: Arturi, Daniel Aarao Reis, et al.
Published: (2025)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs
by: Mushtaq, Erum, et al.
Published: (2025)
by: Mushtaq, Erum, et al.
Published: (2025)
Model Organisms for Emergent Misalignment
by: Turner, Edward, et al.
Published: (2025)
by: Turner, Edward, et al.
Published: (2025)
Convergent Linear Representations of Emergent Misalignment
by: Soligo, Anna, et al.
Published: (2025)
by: Soligo, Anna, et al.
Published: (2025)
The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs
by: Dickson, Craig
Published: (2025)
by: Dickson, Craig
Published: (2025)
In-Training Defenses against Emergent Misalignment in Language Models
by: Kaczér, David, et al.
Published: (2025)
by: Kaczér, David, et al.
Published: (2025)
Understanding Emergent Misalignment via Feature Superposition Geometry
by: Minegishi, Gouki, et al.
Published: (2026)
by: Minegishi, Gouki, et al.
Published: (2026)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
by: Giordani, Jeremiah
Published: (2025)
by: Giordani, Jeremiah
Published: (2025)
Moloch's Bargain: Emergent Misalignment When LLMs Compete for Audiences
by: El, Batu, et al.
Published: (2025)
by: El, Batu, et al.
Published: (2025)
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
by: Chua, James, et al.
Published: (2025)
by: Chua, James, et al.
Published: (2025)
Persona-Model Collapse in Emergent Misalignment
by: Costa, Davi Bastos, et al.
Published: (2026)
by: Costa, Davi Bastos, et al.
Published: (2026)
Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer
by: Askin, Baris, et al.
Published: (2026)
by: Askin, Baris, et al.
Published: (2026)
Neural Total Variation Distance Estimators for Changepoint Detection in News Data
by: Zsolnai, Csaba, et al.
Published: (2025)
by: Zsolnai, Csaba, et al.
Published: (2025)
ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use
by: Tien, Jeremy, et al.
Published: (2026)
by: Tien, Jeremy, et al.
Published: (2026)
Persona Features Control Emergent Misalignment
by: Wang, Miles, et al.
Published: (2025)
by: Wang, Miles, et al.
Published: (2025)
Overtrained, Not Misaligned
by: Schreiber, Joel, et al.
Published: (2026)
by: Schreiber, Joel, et al.
Published: (2026)
Agentic Misalignment: How LLMs Could Be Insider Threats
by: Lynch, Aengus, et al.
Published: (2025)
by: Lynch, Aengus, et al.
Published: (2025)
Decomposed Trust: Privacy, Adversarial Robustness, Ethics, and Fairness in Low-Rank LLMs
by: Asante, Daniel Agyei, et al.
Published: (2025)
by: Asante, Daniel Agyei, et al.
Published: (2025)
GraphTool-Instruction: Revolutionizing Graph Reasoning in LLMs through Decomposed Subtask Instruction
by: Wang, Rongzheng, et al.
Published: (2024)
by: Wang, Rongzheng, et al.
Published: (2024)
BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking
by: Ustaomeroglu, Muhammed, et al.
Published: (2026)
by: Ustaomeroglu, Muhammed, et al.
Published: (2026)
Misalignment from Treating Means as Ends
by: Marklund, Henrik, et al.
Published: (2025)
by: Marklund, Henrik, et al.
Published: (2025)
Unified Parameter-Efficient Unlearning for LLMs
by: Ding, Chenlu, et al.
Published: (2024)
by: Ding, Chenlu, et al.
Published: (2024)
LLM Misalignment via Adversarial RLHF Platforms
by: Entezami, Erfan, et al.
Published: (2025)
by: Entezami, Erfan, et al.
Published: (2025)
Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation
by: Zelikman, Eric, et al.
Published: (2023)
by: Zelikman, Eric, et al.
Published: (2023)
Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
DecompGAIL: Learning Realistic Traffic Behaviors with Decomposed Multi-Agent Generative Adversarial Imitation Learning
by: Guo, Ke, et al.
Published: (2025)
by: Guo, Ke, et al.
Published: (2025)
Indirect Attention: Turning Context Misalignment into a Feature
by: Bahaduri, Bissmella, et al.
Published: (2025)
by: Bahaduri, Bissmella, et al.
Published: (2025)
From Parameters to Behaviors: Unsupervised Compression of the Policy Space
by: Tenedini, Davide, et al.
Published: (2025)
by: Tenedini, Davide, et al.
Published: (2025)
Decomposing Epistemic Uncertainty for Causal Decision Making
by: Rahman, Md Musfiqur, et al.
Published: (2026)
by: Rahman, Md Musfiqur, et al.
Published: (2026)
SDQ: Sparse Decomposed Quantization for LLM Inference
by: Jeong, Geonhwa, et al.
Published: (2024)
by: Jeong, Geonhwa, et al.
Published: (2024)
Decomposing and Editing Predictions by Modeling Model Computation
by: Shah, Harshay, et al.
Published: (2024)
by: Shah, Harshay, et al.
Published: (2024)
Catapult Dynamics and Phase Transitions in Quadratic Nets
by: Meltzer, David, et al.
Published: (2023)
by: Meltzer, David, et al.
Published: (2023)
Zeroth-Order Fine-Tuning of LLMs in Random Subspaces
by: Yu, Ziming, et al.
Published: (2024)
by: Yu, Ziming, et al.
Published: (2024)
Zeroth-Order Fine-Tuning of LLMs with Extreme Sparsity
by: Guo, Wentao, et al.
Published: (2024)
by: Guo, Wentao, et al.
Published: (2024)
Beyond Prior Limits: Addressing Distribution Misalignment in Particle Filtering
by: Shi, Yiwei, et al.
Published: (2025)
by: Shi, Yiwei, et al.
Published: (2025)
Preemptive Detection and Steering of LLM Misalignment via Latent Reachability
by: Karnik, Sathwik, et al.
Published: (2025)
by: Karnik, Sathwik, et al.
Published: (2025)
ConformaDecompose: Explaining Uncertainty via Calibration Localization
by: Yapicioglu, Fatima Rabia, et al.
Published: (2026)
by: Yapicioglu, Fatima Rabia, et al.
Published: (2026)
The Viscosity of Logic: Phase Transitions and Hysteresis in DPO Alignment
by: Pollanen, Marco
Published: (2026)
by: Pollanen, Marco
Published: (2026)
Similar Items
-
Phase Transitions in the Output Distribution of Large Language Models
by: Arnold, Julian, et al.
Published: (2024) -
Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior
by: Arturi, Daniel Aarao Reis, et al.
Published: (2025) -
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025) -
From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs
by: Mushtaq, Erum, et al.
Published: (2025) -
Model Organisms for Emergent Misalignment
by: Turner, Edward, et al.
Published: (2025)