Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Betley, Jan, Tan, Daniel, Warncke, Niels, Sztyber-Betley, Anna, Bao, Xuchan, Soto, Martín, Labenz, Nathan, Evans, Owain |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
par: Dubiński, Jan, et autres
Publié: (2026)
par: Dubiński, Jan, et autres
Publié: (2026)
Tell me about yourself: LLMs are aware of their learned behaviors
par: Betley, Jan, et autres
Publié: (2025)
par: Betley, Jan, et autres
Publié: (2025)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
par: Betley, Jan, et autres
Publié: (2025)
par: Betley, Jan, et autres
Publié: (2025)
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
par: Chua, James, et autres
Publié: (2025)
par: Chua, James, et autres
Publié: (2025)
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
par: Taylor, Mia, et autres
Publié: (2025)
par: Taylor, Mia, et autres
Publié: (2025)
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
par: Cloud, Alex, et autres
Publié: (2025)
par: Cloud, Alex, et autres
Publié: (2025)
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
par: Chua, James, et autres
Publié: (2026)
par: Chua, James, et autres
Publié: (2026)
Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs
par: Afonin, Nikita, et autres
Publié: (2025)
par: Afonin, Nikita, et autres
Publié: (2025)
Emergent misalignment as prompt sensitivity: A research note
par: Wyse, Tim, et autres
Publié: (2025)
par: Wyse, Tim, et autres
Publié: (2025)
Jailbroken Frontier Models Retain Their Capabilities
par: Zhu, Daniel, et autres
Publié: (2026)
par: Zhu, Daniel, et autres
Publié: (2026)
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
par: Hu, Xuhao, et autres
Publié: (2025)
par: Hu, Xuhao, et autres
Publié: (2025)
Persona-Model Collapse in Emergent Misalignment
par: Costa, Davi Bastos, et autres
Publié: (2026)
par: Costa, Davi Bastos, et autres
Publié: (2026)
BadGPT-4o: stripping safety finetuning from GPT models
par: Krupkina, Ekaterina, et autres
Publié: (2024)
par: Krupkina, Ekaterina, et autres
Publié: (2024)
Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Models
par: Panpatil, Siddhant, et autres
Publié: (2025)
par: Panpatil, Siddhant, et autres
Publié: (2025)
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
par: Treutlein, Johannes, et autres
Publié: (2024)
par: Treutlein, Johannes, et autres
Publié: (2024)
Glitch in Time: Exploiting Temporal Misalignment of IMU For Eavesdropping
par: Najeeb, Ahmed, et autres
Publié: (2024)
par: Najeeb, Ahmed, et autres
Publié: (2024)
Agentic Misalignment: How LLMs Could Be Insider Threats
par: Lynch, Aengus, et autres
Publié: (2025)
par: Lynch, Aengus, et autres
Publié: (2025)
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
par: Zhang, Rui, et autres
Publié: (2026)
par: Zhang, Rui, et autres
Publié: (2026)
Narrowing the Gap between TEEs Threat Model and Deployment Strategies
par: Rezabek, Filip, et autres
Publié: (2025)
par: Rezabek, Filip, et autres
Publié: (2025)
Badllama 3: removing safety finetuning from Llama 3 in minutes
par: Volkov, Dmitrii
Publié: (2024)
par: Volkov, Dmitrii
Publié: (2024)
Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures
par: Su, Yanghao, et autres
Publié: (2026)
par: Su, Yanghao, et autres
Publié: (2026)
Emergent (In)Security of Multi-Cloud Environments
par: Reece, Morgan, et autres
Publié: (2023)
par: Reece, Morgan, et autres
Publié: (2023)
Security Analysis of Time-of-Arrival Estimation via Cross-Correlation under Narrow-Band Conditions
par: Anliker, Claudio, et autres
Publié: (2026)
par: Anliker, Claudio, et autres
Publié: (2026)
LASHED: LLMs And Static Hardware Analysis for Early Detection of RTL Bugs
par: Ahmad, Baleegh, et autres
Publié: (2025)
par: Ahmad, Baleegh, et autres
Publié: (2025)
Investigating the Effect of Misalignment on Membership Privacy in the White-box Setting
par: Cretu, Ana-Maria, et autres
Publié: (2023)
par: Cretu, Ana-Maria, et autres
Publié: (2023)
D4+: Emergent Adversarial Driving Maneuvers with Approximate Functional Optimization
par: Barbosa, Diego Ortiz, et autres
Publié: (2025)
par: Barbosa, Diego Ortiz, et autres
Publié: (2025)
Narrow Secret Loyalty Dodges Black-Box Audits
par: Lamerton, Alfie, et autres
Publié: (2026)
par: Lamerton, Alfie, et autres
Publié: (2026)
Zebrafix: Mitigating Memory-Centric Side-Channel Leakage via Interleaving
par: Pätschke, Anna, et autres
Publié: (2025)
par: Pätschke, Anna, et autres
Publié: (2025)
Evaluating Google's Protected Audience Protocol
par: Long, Minjun, et autres
Publié: (2024)
par: Long, Minjun, et autres
Publié: (2024)
Composable Attestation: A Generalized Framework for Continuous and Incremental Trust in AI-Driven Distributed Systems
par: Sun, Sheng, et autres
Publié: (2026)
par: Sun, Sheng, et autres
Publié: (2026)
I can't recognize (yet): Delayed Rendering to Defeat Visual Phishing Detectors
par: Yuan, Ying, et autres
Publié: (2026)
par: Yuan, Ying, et autres
Publié: (2026)
Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models
par: Ni, Zhenyang, et autres
Publié: (2024)
par: Ni, Zhenyang, et autres
Publié: (2024)
Bots can Snoop: Uncovering and Mitigating Privacy Risks of Bots in Group Chats
par: Chou, Kai-Hsiang, et autres
Publié: (2024)
par: Chou, Kai-Hsiang, et autres
Publié: (2024)
XBreaking: Understanding how LLMs security alignment can be broken
par: Arazzi, Marco, et autres
Publié: (2025)
par: Arazzi, Marco, et autres
Publié: (2025)
DP-RuL: Differentially-Private Rule Learning for Clinical Decision Support Systems
par: Lamp, Josephine, et autres
Publié: (2024)
par: Lamp, Josephine, et autres
Publié: (2024)
Building an Open Source Operational Technology Pentesting Platform: Lessons from LINICS
par: Rashid, Awais, et autres
Publié: (2026)
par: Rashid, Awais, et autres
Publié: (2026)
Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment Design
par: Eshuijs, Leon, et autres
Publié: (2026)
par: Eshuijs, Leon, et autres
Publié: (2026)
Adversarial Examples are Misaligned in Diffusion Model Manifolds
par: Lorenz, Peter, et autres
Publié: (2024)
par: Lorenz, Peter, et autres
Publié: (2024)
Spoofing attack augmentation: can differently-trained attack models improve generalisation?
par: Ge, Wanying, et autres
Publié: (2023)
par: Ge, Wanying, et autres
Publié: (2023)
Obelix: Mitigating Side-Channels Through Dynamic Obfuscation
par: Wichelmann, Jan, et autres
Publié: (2025)
par: Wichelmann, Jan, et autres
Publié: (2025)
Documents similaires
-
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
par: Dubiński, Jan, et autres
Publié: (2026) -
Tell me about yourself: LLMs are aware of their learned behaviors
par: Betley, Jan, et autres
Publié: (2025) -
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
par: Betley, Jan, et autres
Publié: (2025) -
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
par: Chua, James, et autres
Publié: (2025) -
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
par: Taylor, Mia, et autres
Publié: (2025)