Overtrained, Not Misaligned
Fuente:
arXiv
Saved in:
| Main Authors: | Schreiber, Joel, Goldstein, Ariel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Model Organisms for Emergent Misalignment
by: Turner, Edward, et al.
Published: (2025)
by: Turner, Edward, et al.
Published: (2025)
Convergent Linear Representations of Emergent Misalignment
by: Soligo, Anna, et al.
Published: (2025)
by: Soligo, Anna, et al.
Published: (2025)
Misalignment from Treating Means as Ends
by: Marklund, Henrik, et al.
Published: (2025)
by: Marklund, Henrik, et al.
Published: (2025)
LLM Misalignment via Adversarial RLHF Platforms
by: Entezami, Erfan, et al.
Published: (2025)
by: Entezami, Erfan, et al.
Published: (2025)
Understanding Emergent Misalignment via Feature Superposition Geometry
by: Minegishi, Gouki, et al.
Published: (2026)
by: Minegishi, Gouki, et al.
Published: (2026)
Indirect Attention: Turning Context Misalignment into a Feature
by: Bahaduri, Bissmella, et al.
Published: (2025)
by: Bahaduri, Bissmella, et al.
Published: (2025)
In-Training Defenses against Emergent Misalignment in Language Models
by: Kaczér, David, et al.
Published: (2025)
by: Kaczér, David, et al.
Published: (2025)
Beyond Prior Limits: Addressing Distribution Misalignment in Particle Filtering
by: Shi, Yiwei, et al.
Published: (2025)
by: Shi, Yiwei, et al.
Published: (2025)
Preemptive Detection and Steering of LLM Misalignment via Latent Reachability
by: Karnik, Sathwik, et al.
Published: (2025)
by: Karnik, Sathwik, et al.
Published: (2025)
Quantifying the Influences on Probabilistic Wind Power Forecasts
by: Schreiber, Jens, et al.
Published: (2018)
by: Schreiber, Jens, et al.
Published: (2018)
ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use
by: Tien, Jeremy, et al.
Published: (2026)
by: Tien, Jeremy, et al.
Published: (2026)
Distance-Misaligned Training in Graph Transformers and Adaptive Graph-Aware Control
by: Hou, Qinhan, et al.
Published: (2026)
by: Hou, Qinhan, et al.
Published: (2026)
Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior
by: Arturi, Daniel Aarao Reis, et al.
Published: (2025)
by: Arturi, Daniel Aarao Reis, et al.
Published: (2025)
Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment
by: Arnold, Julian, et al.
Published: (2025)
by: Arnold, Julian, et al.
Published: (2025)
From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs
by: Mushtaq, Erum, et al.
Published: (2025)
by: Mushtaq, Erum, et al.
Published: (2025)
A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns
by: Yehudai, Asaf, et al.
Published: (2024)
by: Yehudai, Asaf, et al.
Published: (2024)
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
by: Liang, Kaiqu, et al.
Published: (2025)
by: Liang, Kaiqu, et al.
Published: (2025)
Silent Inconsistency in Data-Parallel Full Fine-Tuning: Diagnosing Worker-Level Optimization Misalignment
by: Li, Hong, et al.
Published: (2026)
by: Li, Hong, et al.
Published: (2026)
Deep Learning for Optical Misalignment Diagnostics in Multi-Lens Imaging Systems
by: Slor, Tomer, et al.
Published: (2025)
by: Slor, Tomer, et al.
Published: (2025)
Epistemic Traps: Rational Misalignment Driven by Model Misspecification
by: Xu, Xingcheng, et al.
Published: (2026)
by: Xu, Xingcheng, et al.
Published: (2026)
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
by: Chua, James, et al.
Published: (2025)
by: Chua, James, et al.
Published: (2025)
Agentic Misalignment: How LLMs Could Be Insider Threats
by: Lynch, Aengus, et al.
Published: (2025)
by: Lynch, Aengus, et al.
Published: (2025)
Ulterior Motives: Detecting Misaligned Reasoning in Continuous Thought Models
by: Ramjee, Sharan
Published: (2026)
by: Ramjee, Sharan
Published: (2026)
Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer
by: Askin, Baris, et al.
Published: (2026)
by: Askin, Baris, et al.
Published: (2026)
Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation
by: Hahm, Dongyoon, et al.
Published: (2025)
by: Hahm, Dongyoon, et al.
Published: (2025)
PASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual Learning
by: Hou, Zhiyan, et al.
Published: (2026)
by: Hou, Zhiyan, et al.
Published: (2026)
The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs
by: Dickson, Craig
Published: (2025)
by: Dickson, Craig
Published: (2025)
Persona-Model Collapse in Emergent Misalignment
by: Costa, Davi Bastos, et al.
Published: (2026)
by: Costa, Davi Bastos, et al.
Published: (2026)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
by: Giordani, Jeremiah
Published: (2025)
by: Giordani, Jeremiah
Published: (2025)
Speculating Experts Accelerates Inference for Mixture-of-Experts
by: Madan, Vivan, et al.
Published: (2026)
by: Madan, Vivan, et al.
Published: (2026)
MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness
by: Hathidara, Ashutosh, et al.
Published: (2026)
by: Hathidara, Ashutosh, et al.
Published: (2026)
Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
by: Hahm, Dongyoon, et al.
Published: (2026)
by: Hahm, Dongyoon, et al.
Published: (2026)
CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval
by: Senthil, Vaishali, et al.
Published: (2026)
by: Senthil, Vaishali, et al.
Published: (2026)
Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky
by: Hathidara, Ashutosh, et al.
Published: (2025)
by: Hathidara, Ashutosh, et al.
Published: (2025)
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
by: Williams, Kai, et al.
Published: (2025)
by: Williams, Kai, et al.
Published: (2025)
Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
Does Visual Rendering Bypass Tokenization? Investigating Script-Tokenizer Misalignment in Pixel-Based Language Models
by: Susanto, Lucky, et al.
Published: (2026)
by: Susanto, Lucky, et al.
Published: (2026)
Can We Predict Alignment Before Models Finish Thinking? Towards Monitoring Misaligned Reasoning Models
by: Chan, Yik Siu, et al.
Published: (2025)
by: Chan, Yik Siu, et al.
Published: (2025)
Zero-shot Multivariate Time Series Forecasting Using Tabular Prior Fitted Networks
by: Jayawardhana, Mayuka, et al.
Published: (2026)
by: Jayawardhana, Mayuka, et al.
Published: (2026)
Similar Items
-
Model Organisms for Emergent Misalignment
by: Turner, Edward, et al.
Published: (2025) -
Convergent Linear Representations of Emergent Misalignment
by: Soligo, Anna, et al.
Published: (2025) -
Misalignment from Treating Means as Ends
by: Marklund, Henrik, et al.
Published: (2025) -
LLM Misalignment via Adversarial RLHF Platforms
by: Entezami, Erfan, et al.
Published: (2025) -
Understanding Emergent Misalignment via Feature Superposition Geometry
by: Minegishi, Gouki, et al.
Published: (2026)