Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
Fuente:
arXiv
Saved in:
| Main Authors: | Dubiński, Jan, Betley, Jan, Sztyber-Betley, Anna, Tan, Daniel, Evans, Owain |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Tell me about yourself: LLMs are aware of their learned behaviors
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
by: Cloud, Alex, et al.
Published: (2025)
by: Cloud, Alex, et al.
Published: (2025)
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
by: Taylor, Mia, et al.
Published: (2025)
by: Taylor, Mia, et al.
Published: (2025)
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
by: Chua, James, et al.
Published: (2025)
by: Chua, James, et al.
Published: (2025)
Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
by: Pawlak, Stanisław, et al.
Published: (2025)
by: Pawlak, Stanisław, et al.
Published: (2025)
LLMs can hide text in other text of the same length
by: Norelli, Antonio, et al.
Published: (2025)
by: Norelli, Antonio, et al.
Published: (2025)
Emergent misalignment as prompt sensitivity: A research note
by: Wyse, Tim, et al.
Published: (2025)
by: Wyse, Tim, et al.
Published: (2025)
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
by: Treutlein, Johannes, et al.
Published: (2024)
by: Treutlein, Johannes, et al.
Published: (2024)
Efficient LLM Moderation with Multi-Layer Latent Prototypes
by: Chrabąszcz, Maciej, et al.
Published: (2025)
by: Chrabąszcz, Maciej, et al.
Published: (2025)
I can't see it but I can Fine-tune it: On Encrypted Fine-tuning of Transformers using Fully Homomorphic Encryption
by: Panzade, Prajwal, et al.
Published: (2024)
by: Panzade, Prajwal, et al.
Published: (2024)
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
by: Chua, James, et al.
Published: (2026)
by: Chua, James, et al.
Published: (2026)
SoK: Verifiable Cross-Silo FL
by: Korneev, Aleksei, et al.
Published: (2024)
by: Korneev, Aleksei, et al.
Published: (2024)
RMF: A Risk Measurement Framework for Machine Learning Models
by: Schröder, Jan, et al.
Published: (2024)
by: Schröder, Jan, et al.
Published: (2024)
CDI: Copyrighted Data Identification in Diffusion Models
by: Dubiński, Jan, et al.
Published: (2024)
by: Dubiński, Jan, et al.
Published: (2024)
Sparsity in neural networks can improve their privacy
by: Gonon, Antoine, et al.
Published: (2023)
by: Gonon, Antoine, et al.
Published: (2023)
XBreaking: Understanding how LLMs security alignment can be broken
by: Arazzi, Marco, et al.
Published: (2025)
by: Arazzi, Marco, et al.
Published: (2025)
A Multiparty Homomorphic Encryption Approach to Confidential Federated Kaplan Meier Survival Analysis
by: Veeraragavan, Narasimha Raghavan, et al.
Published: (2024)
by: Veeraragavan, Narasimha Raghavan, et al.
Published: (2024)
Do Parameters Reveal More than Loss for Membership Inference?
by: Suri, Anshuman, et al.
Published: (2024)
by: Suri, Anshuman, et al.
Published: (2024)
Dissecting Distribution Inference
by: Suri, Anshuman, et al.
Published: (2022)
by: Suri, Anshuman, et al.
Published: (2022)
Prompt Injection Attacks on Large Language Models in Oncology
by: Clusmann, Jan, et al.
Published: (2024)
by: Clusmann, Jan, et al.
Published: (2024)
Protect and Extend -- Using GANs for Synthetic Data Generation of Time-Series Medical Records
by: Ashrafi, Navid, et al.
Published: (2024)
by: Ashrafi, Navid, et al.
Published: (2024)
A Visualized Malware Detection Framework with CNN and Conditional GAN
by: Wang, Fang, et al.
Published: (2024)
by: Wang, Fang, et al.
Published: (2024)
Conditional Adversarial Fragility in Financial Machine Learning under Macroeconomic Stress
by: Baviskar, Samruddhi
Published: (2025)
by: Baviskar, Samruddhi
Published: (2025)
Exploring Feature Importance and Explainability Towards Enhanced ML-Based DoS Detection in AI Systems
by: Yakubu, Paul Badu, et al.
Published: (2024)
by: Yakubu, Paul Badu, et al.
Published: (2024)
Quantum-Augmented AI/ML for O-RAN: Hierarchical Threat Detection with Synergistic Intelligence and Interpretability (Technical Report)
by: Le, Tan, et al.
Published: (2025)
by: Le, Tan, et al.
Published: (2025)
Detecting Malicious AI Agents Through Simulated Interactions
by: Pi, Yulu, et al.
Published: (2025)
by: Pi, Yulu, et al.
Published: (2025)
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
by: Lin, Shuyi, et al.
Published: (2025)
by: Lin, Shuyi, et al.
Published: (2025)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
by: Laine, Rudolf, et al.
Published: (2024)
by: Laine, Rudolf, et al.
Published: (2024)
On Stealing Graph Neural Network Models
by: Podhajski, Marcin, et al.
Published: (2025)
by: Podhajski, Marcin, et al.
Published: (2025)
Unlink to Unlearn: Simplifying Edge Unlearning in GNNs
by: Tan, Jiajun, et al.
Published: (2024)
by: Tan, Jiajun, et al.
Published: (2024)
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
by: Braun, Tobias, et al.
Published: (2025)
by: Braun, Tobias, et al.
Published: (2025)
AED: Automatic Discovery of Effective and Diverse Vulnerabilities for Autonomous Driving Policy with Large Language Models
by: Qiu, Le, et al.
Published: (2025)
by: Qiu, Le, et al.
Published: (2025)
Analyzing Consumer IoT Traffic from Security and Privacy Perspectives: a Comprehensive Survey
by: Jia, Yan, et al.
Published: (2024)
by: Jia, Yan, et al.
Published: (2024)
Incentivising the federation: gradient-based metrics for data selection and valuation in private decentralised training
by: Usynin, Dmitrii, et al.
Published: (2023)
by: Usynin, Dmitrii, et al.
Published: (2023)
Privacy Preserving Federated Learning with Convolutional Variational Bottlenecks
by: Scheliga, Daniel, et al.
Published: (2023)
by: Scheliga, Daniel, et al.
Published: (2023)
Combining Stochastic Defenses to Resist Gradient Inversion: An Ablation Study
by: Scheliga, Daniel, et al.
Published: (2022)
by: Scheliga, Daniel, et al.
Published: (2022)
Jailbroken Frontier Models Retain Their Capabilities
by: Zhu, Daniel, et al.
Published: (2026)
by: Zhu, Daniel, et al.
Published: (2026)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
by: Lin, Shi, et al.
Published: (2024)
by: Lin, Shi, et al.
Published: (2024)
Similar Items
-
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025) -
Tell me about yourself: LLMs are aware of their learned behaviors
by: Betley, Jan, et al.
Published: (2025) -
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
by: Betley, Jan, et al.
Published: (2025) -
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
by: Cloud, Alex, et al.
Published: (2025) -
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
by: Taylor, Mia, et al.
Published: (2025)