LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Xuhao, Wang, Peng, Lu, Xiaoya, Liu, Dongrui, Huang, Xuanjing, Shao, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLSBench: Unveiling Visual Leakage in Multimodal Safety
by: Hu, Xuhao, et al.
Published: (2024)
by: Hu, Xuhao, et al.
Published: (2024)
Persona-Model Collapse in Emergent Misalignment
by: Costa, Davi Bastos, et al.
Published: (2026)
by: Costa, Davi Bastos, et al.
Published: (2026)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Models
by: Panpatil, Siddhant, et al.
Published: (2025)
by: Panpatil, Siddhant, et al.
Published: (2025)
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
by: Zhang, Rui, et al.
Published: (2026)
by: Zhang, Rui, et al.
Published: (2026)
X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability
by: Lu, Xiaoya, et al.
Published: (2025)
by: Lu, Xiaoya, et al.
Published: (2025)
Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures
by: Su, Yanghao, et al.
Published: (2026)
by: Su, Yanghao, et al.
Published: (2026)
RvB: Automating AI System Hardening via Iterative Red-Blue Games
by: Huang, Lige, et al.
Published: (2026)
by: Huang, Lige, et al.
Published: (2026)
T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation
by: Li, Lijun, et al.
Published: (2025)
by: Li, Lijun, et al.
Published: (2025)
Glitch in Time: Exploiting Temporal Misalignment of IMU For Eavesdropping
by: Najeeb, Ahmed, et al.
Published: (2024)
by: Najeeb, Ahmed, et al.
Published: (2024)
Agentic Misalignment: How LLMs Could Be Insider Threats
by: Lynch, Aengus, et al.
Published: (2025)
by: Lynch, Aengus, et al.
Published: (2025)
Malicious and Unintentional Disclosure Risks in Large Language Models for Code Generation
by: Rabin, Rafiqul, et al.
Published: (2025)
by: Rabin, Rafiqul, et al.
Published: (2025)
No Attacker Needed: Unintentional Cross-User Contamination in Shared-State LLM Agents
by: Yang, Tiankai, et al.
Published: (2026)
by: Yang, Tiankai, et al.
Published: (2026)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
REEF: Representation Encoding Fingerprints for Large Language Models
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
by: Liu, Zicheng, et al.
Published: (2025)
by: Liu, Zicheng, et al.
Published: (2025)
Investigating the Effect of Misalignment on Membership Privacy in the White-box Setting
by: Cretu, Ana-Maria, et al.
Published: (2023)
by: Cretu, Ana-Maria, et al.
Published: (2023)
Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
Adversarial Examples are Misaligned in Diffusion Model Manifolds
by: Lorenz, Peter, et al.
Published: (2024)
by: Lorenz, Peter, et al.
Published: (2024)
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
by: Dassanayake, Rishane, et al.
Published: (2025)
by: Dassanayake, Rishane, et al.
Published: (2025)
DeepSight: An All-in-One LM Safety Toolkit
by: Zhang, Bo, et al.
Published: (2026)
by: Zhang, Bo, et al.
Published: (2026)
Fingerprinting LLMs via Prompt Injection
by: Hu, Yuepeng, et al.
Published: (2025)
by: Hu, Yuepeng, et al.
Published: (2025)
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
by: Li, Lijun, et al.
Published: (2024)
by: Li, Lijun, et al.
Published: (2024)
Separator Injection Attack: Uncovering Dialogue Biases in Large Language Models Caused by Role Separators
by: Li, Xitao, et al.
Published: (2025)
by: Li, Xitao, et al.
Published: (2025)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
by: Zhang, Chiyu, et al.
Published: (2025)
by: Zhang, Chiyu, et al.
Published: (2025)
HLPD: Aligning LLMs to Human Language Preference for Machine-Revised Text Detection
by: Dai, Fangqi, et al.
Published: (2025)
by: Dai, Fangqi, et al.
Published: (2025)
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
by: Long, Zhuohang, et al.
Published: (2025)
by: Long, Zhuohang, et al.
Published: (2025)
A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness
by: Luo, Xuan, et al.
Published: (2025)
by: Luo, Xuan, et al.
Published: (2025)
Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment Design
by: Eshuijs, Leon, et al.
Published: (2026)
by: Eshuijs, Leon, et al.
Published: (2026)
Subspace Defense: Discarding Adversarial Perturbations by Learning a Subspace for Clean Signals
by: Zheng, Rui, et al.
Published: (2024)
by: Zheng, Rui, et al.
Published: (2024)
The Ethics of Interaction: Mitigating Security Threats in LLMs
by: Kumar, Ashutosh, et al.
Published: (2024)
by: Kumar, Ashutosh, et al.
Published: (2024)
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
by: Williams, Kai, et al.
Published: (2025)
by: Williams, Kai, et al.
Published: (2025)
Seeing is Deceiving: Mirror-Based LiDAR Spoofing for Autonomous Vehicle Deception
by: Yahia, Selma, et al.
Published: (2025)
by: Yahia, Selma, et al.
Published: (2025)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
by: Huang, Caishuang, et al.
Published: (2024)
by: Huang, Caishuang, et al.
Published: (2024)
Enabling Efficient Attack Investigation via Human-in-the-Loop Security Analysis
by: Tsegai, Saimon Amanuel, et al.
Published: (2022)
by: Tsegai, Saimon Amanuel, et al.
Published: (2022)
Doubly-Universal Adversarial Perturbations: Deceiving Vision-Language Models Across Both Images and Text with a Single Perturbation
by: Kim, Hee-Seon, et al.
Published: (2024)
by: Kim, Hee-Seon, et al.
Published: (2024)
FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity
by: Cui, Shiyao, et al.
Published: (2023)
by: Cui, Shiyao, et al.
Published: (2023)
A Human Study of Cognitive Biases in Web Application Security
by: Yang, Yuwei, et al.
Published: (2025)
by: Yang, Yuwei, et al.
Published: (2025)
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
by: Liu, Yepeng, et al.
Published: (2025)
by: Liu, Yepeng, et al.
Published: (2025)
Ingest-And-Ground: Dispelling Hallucinations from Continually-Pretrained LLMs with RAG
by: Fang, Chenhao, et al.
Published: (2024)
by: Fang, Chenhao, et al.
Published: (2024)
Similar Items
-
VLSBench: Unveiling Visual Leakage in Multimodal Safety
by: Hu, Xuhao, et al.
Published: (2024) -
Persona-Model Collapse in Emergent Misalignment
by: Costa, Davi Bastos, et al.
Published: (2026) -
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025) -
Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Models
by: Panpatil, Siddhant, et al.
Published: (2025) -
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
by: Zhang, Rui, et al.
Published: (2026)