Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer
Fuente:
arXiv
Saved in:
| Main Authors: | Askin, Baris, Ustaomeroglu, Muhammed, Nayak, Anupam, Joshi, Gauri, Qu, Guannan, Joe-Wong, Carlee |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
by: Ustaomeroglu, Muhammed, et al.
Published: (2025)
by: Ustaomeroglu, Muhammed, et al.
Published: (2025)
PubSwap: Public-Data Off-Policy Coordination for Federated RLVR
by: Nayak, Anupam, et al.
Published: (2026)
by: Nayak, Anupam, et al.
Published: (2026)
BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking
by: Ustaomeroglu, Muhammed, et al.
Published: (2026)
by: Ustaomeroglu, Muhammed, et al.
Published: (2026)
Towards Effective Theory of LLMs: A Representation Learning Approach
by: Ustaomeroglu, Muhammed, et al.
Published: (2026)
by: Ustaomeroglu, Muhammed, et al.
Published: (2026)
Federated Communication-Efficient Multi-Objective Optimization
by: Askin, Baris, et al.
Published: (2024)
by: Askin, Baris, et al.
Published: (2024)
FedAST: Federated Asynchronous Simultaneous Training
by: Askin, Baris, et al.
Published: (2024)
by: Askin, Baris, et al.
Published: (2024)
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization
by: Ustaomeroglu, Muhammed, et al.
Published: (2025)
by: Ustaomeroglu, Muhammed, et al.
Published: (2025)
Federate the Router: Learning Language Model Routers with Sparse and Decentralized Evaluations
by: Askin, Baris, et al.
Published: (2026)
by: Askin, Baris, et al.
Published: (2026)
Ravan: Multi-Head Low-Rank Adaptation for Federated Fine-Tuning
by: Raje, Arian, et al.
Published: (2025)
by: Raje, Arian, et al.
Published: (2025)
Not All Turns Are Equally Hard: Adaptive Thinking Budgets For Efficient Multi-Turn Reasoning
by: Jali, Neharika, et al.
Published: (2026)
by: Jali, Neharika, et al.
Published: (2026)
MELINOE: Fine-Tuning Enables Memory-Efficient Inference for Mixture-of-Experts Models
by: Raje, Arian, et al.
Published: (2026)
by: Raje, Arian, et al.
Published: (2026)
Reviving Stale Updates: Data-Free Knowledge Distillation for Asynchronous Federated Learning
by: Askin, Baris, et al.
Published: (2025)
by: Askin, Baris, et al.
Published: (2025)
QMOS: Enhancing LLMs for Telecommunication with Question Masked loss and Option Shuffling
by: Guda, Blessed, et al.
Published: (2024)
by: Guda, Blessed, et al.
Published: (2024)
DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent Collaboration
by: Nourzad, Narjes, et al.
Published: (2025)
by: Nourzad, Narjes, et al.
Published: (2025)
Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
by: Aden-Ali, Ishaq, et al.
Published: (2026)
by: Aden-Ali, Ishaq, et al.
Published: (2026)
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
by: Chua, James, et al.
Published: (2025)
by: Chua, James, et al.
Published: (2025)
Persona-Model Collapse in Emergent Misalignment
by: Costa, Davi Bastos, et al.
Published: (2026)
by: Costa, Davi Bastos, et al.
Published: (2026)
MIRA: Memory-Integrated Reinforcement Learning Agent with Limited LLM Guidance
by: Nourzad, Narjes, et al.
Published: (2026)
by: Nourzad, Narjes, et al.
Published: (2026)
Neural Combinatorial Clustered Bandits for Recommendation Systems
by: Atalar, Baran, et al.
Published: (2024)
by: Atalar, Baran, et al.
Published: (2024)
Memory-Based Advantage Shaping for LLM-Guided Reinforcement Learning
by: Nourzad, Narjes, et al.
Published: (2026)
by: Nourzad, Narjes, et al.
Published: (2026)
The Five Ws of Multi-Agent Communication: Who Talks to Whom, When, What, and Why -- A Survey from MARL to Emergent Language and LLMs
by: Chen, Jingdi, et al.
Published: (2026)
by: Chen, Jingdi, et al.
Published: (2026)
Transformer-Based Scalable Multi-Agent Reinforcement Learning for Networked Systems with Long-Range Interactions
by: Sinha, Vidur, et al.
Published: (2025)
by: Sinha, Vidur, et al.
Published: (2025)
The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs
by: Dickson, Craig
Published: (2025)
by: Dickson, Craig
Published: (2025)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
by: Soligo, Anna, et al.
Published: (2026)
by: Soligo, Anna, et al.
Published: (2026)
Efficient Contextual LLM Cascades through Budget-Constrained Policy Learning
by: Zhang, Xuechen, et al.
Published: (2024)
by: Zhang, Xuechen, et al.
Published: (2024)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
by: Giordani, Jeremiah
Published: (2025)
by: Giordani, Jeremiah
Published: (2025)
Matrix-Space Reinforcement Learning for Reusing Local Transition Geometry
by: Zhang, Zuyuan, et al.
Published: (2026)
by: Zhang, Zuyuan, et al.
Published: (2026)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Reframing Data Value for Large Language Models Through the Lens of Plausibility
by: Rammal, Mohamad Rida, et al.
Published: (2024)
by: Rammal, Mohamad Rida, et al.
Published: (2024)
Quantifying and Mitigating Selection Bias in LLMs: A Transferable LoRA Fine-Tuning and Efficient Majority Voting Approach
by: Guda, Blessed, et al.
Published: (2025)
by: Guda, Blessed, et al.
Published: (2025)
Epistemic Traps: Rational Misalignment Driven by Model Misspecification
by: Xu, Xingcheng, et al.
Published: (2026)
by: Xu, Xingcheng, et al.
Published: (2026)
An LLM-Based Digital Twin for Optimizing Human-in-the Loop Systems
by: Yang, Hanqing, et al.
Published: (2024)
by: Yang, Hanqing, et al.
Published: (2024)
FedTLU: Federated Learning with Targeted Layer Updates
by: Park, Jong-Ik, et al.
Published: (2024)
by: Park, Jong-Ik, et al.
Published: (2024)
Federated Learning with Flexible Architectures
by: Park, Jong-Ik, et al.
Published: (2024)
by: Park, Jong-Ik, et al.
Published: (2024)
CoRAST: Towards Foundation Model-Powered Correlated Data Analysis in Resource-Constrained CPS and IoT
by: Hu, Yi, et al.
Published: (2024)
by: Hu, Yi, et al.
Published: (2024)
Named Entity Recognition for Payment Data Using NLP
by: Nayak, Srikumar
Published: (2026)
by: Nayak, Srikumar
Published: (2026)
Model Organisms for Emergent Misalignment
by: Turner, Edward, et al.
Published: (2025)
by: Turner, Edward, et al.
Published: (2025)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
by: Schrodi, Simon, et al.
Published: (2025)
by: Schrodi, Simon, et al.
Published: (2025)
Semantic Containment as a Fundamental Property of Emergent Misalignment
by: Saxena, Rohan
Published: (2026)
by: Saxena, Rohan
Published: (2026)
Efficient Prompt Optimization Through the Lens of Best Arm Identification
by: Shi, Chengshuai, et al.
Published: (2024)
by: Shi, Chengshuai, et al.
Published: (2024)
Similar Items
-
Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
by: Ustaomeroglu, Muhammed, et al.
Published: (2025) -
PubSwap: Public-Data Off-Policy Coordination for Federated RLVR
by: Nayak, Anupam, et al.
Published: (2026) -
BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking
by: Ustaomeroglu, Muhammed, et al.
Published: (2026) -
Towards Effective Theory of LLMs: A Representation Learning Approach
by: Ustaomeroglu, Muhammed, et al.
Published: (2026) -
Federated Communication-Efficient Multi-Objective Optimization
by: Askin, Baris, et al.
Published: (2024)