The Shadow Self: Intrinsic Value Misalignment in Large Language Model Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Chen, Il, Kim Young, Yang, Yuan, Su, Wenhao, Zhang, Yilin, Gong, Xueluan, Wang, Qian, Zheng, Yongsen, Liu, Ziyao, Lam, Kwok-Yan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
Neutralizing Backdoors through Information Conflicts for Large Language Models
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
Threats, Attacks, and Defenses in Machine Unlearning: A Survey
by: Liu, Ziyao, et al.
Published: (2024)
by: Liu, Ziyao, et al.
Published: (2024)
Attribution Techniques for Mitigating Hallucinated Information in RAG Systems: A Survey
by: Zhao, Yuqing, et al.
Published: (2026)
by: Zhao, Yuqing, et al.
Published: (2026)
Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agents
by: Zhou, Kaiyu, et al.
Published: (2026)
by: Zhou, Kaiyu, et al.
Published: (2026)
PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs
by: Gong, Xueluan, et al.
Published: (2024)
by: Gong, Xueluan, et al.
Published: (2024)
Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems
by: Gao, Jiaxin, et al.
Published: (2025)
by: Gao, Jiaxin, et al.
Published: (2025)
FROC: A Unified Framework with Risk-Optimized Control for Machine Unlearning in LLMs
by: Goh, Si Qi, et al.
Published: (2025)
by: Goh, Si Qi, et al.
Published: (2025)
Why Multi-Interest Fairness Matters: Hypergraph Contrastive Multi-Interest Learning for Fair Conversational Recommender System
by: Zheng, Yongsen, et al.
Published: (2025)
by: Zheng, Yongsen, et al.
Published: (2025)
Plato's Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
Hidden Data Privacy Breaches in Federated Learning
by: Gong, Xueluan, et al.
Published: (2024)
by: Gong, Xueluan, et al.
Published: (2024)
An Investigation into Value Misalignment in LLM-Generated Texts for Cultural Heritage
by: Bu, Fan, et al.
Published: (2025)
by: Bu, Fan, et al.
Published: (2025)
A Survey on Facial Image Privacy Preservation in Cloud-Based Services
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
by: Li, Songze, et al.
Published: (2025)
by: Li, Songze, et al.
Published: (2025)
Process-of-Thought Reasoning for Videos
by: Zhang, Jusheng, et al.
Published: (2026)
by: Zhang, Jusheng, et al.
Published: (2026)
Efficient Privacy-Preserving Retrieval Augmented Generation with Distance-Preserving Encryption
by: Ye, Huanyi, et al.
Published: (2026)
by: Ye, Huanyi, et al.
Published: (2026)
An Effective and Resilient Backdoor Attack Framework against Deep Neural Networks and Vision Transformers
by: Gong, Xueluan, et al.
Published: (2024)
by: Gong, Xueluan, et al.
Published: (2024)
Confidence Matters: Revisiting Intrinsic Self-Correction Capabilities of Large Language Models
by: Li, Loka, et al.
Published: (2024)
by: Li, Loka, et al.
Published: (2024)
Megatron: Evasive Clean-Label Backdoor Attacks against Vision Transformer
by: Gong, Xueluan, et al.
Published: (2024)
by: Gong, Xueluan, et al.
Published: (2024)
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
by: Naik, Akshat, et al.
Published: (2025)
by: Naik, Akshat, et al.
Published: (2025)
Spectral Gating Networks
by: Zhang, Jusheng, et al.
Published: (2026)
by: Zhang, Jusheng, et al.
Published: (2026)
Towards Efficient and Certified Recovery from Poisoning Attacks in Federated Learning
by: Jiang, Yu, et al.
Published: (2024)
by: Jiang, Yu, et al.
Published: (2024)
Understanding Hardness of Vision-Language Compositionality from A Token-level Causal Lens
by: Chen, Ziliang, et al.
Published: (2025)
by: Chen, Ziliang, et al.
Published: (2025)
A Survey on Federated Unlearning: Challenges, Methods, and Future Directions
by: Liu, Ziyao, et al.
Published: (2023)
by: Liu, Ziyao, et al.
Published: (2023)
Guaranteeing Data Privacy in Federated Unlearning with Dynamic User Participation
by: Liu, Ziyao, et al.
Published: (2024)
by: Liu, Ziyao, et al.
Published: (2024)
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
Weighted SurvClipper : Nonlinear Prognostic Biomarker Selection Incorporating Historical Information for Survival Risk With Controlled FDR
by: Yaxian Chen, et al.
Published: (2025)
by: Yaxian Chen, et al.
Published: (2025)
TrojanDam: Detection-Free Backdoor Defense in Federated Learning through Proactive Model Robustification utilizing OOD Data
by: Dai, Yanbo, et al.
Published: (2025)
by: Dai, Yanbo, et al.
Published: (2025)
Cross-modal Causal Relation Alignment for Video Question Grounding
by: Chen, Weixing, et al.
Published: (2025)
by: Chen, Weixing, et al.
Published: (2025)
Proactive Detection of Physical Inter-rule Vulnerabilities in IoT Services Using a Deep Learning Approach
by: Huang, Bing, et al.
Published: (2024)
by: Huang, Bing, et al.
Published: (2024)
Efficient Federated Unlearning with Adaptive Differential Privacy Preservation
by: Jiang, Yu, et al.
Published: (2024)
by: Jiang, Yu, et al.
Published: (2024)
Certifying the Right to Be Forgotten: Primal-Dual Optimization for Sample and Label Unlearning in Vertical Federated Learning
by: Jiang, Yu, et al.
Published: (2025)
by: Jiang, Yu, et al.
Published: (2025)
From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents
by: Wang, Xinyue, et al.
Published: (2026)
by: Wang, Xinyue, et al.
Published: (2026)
ARMOR: Shielding Unlearnable Examples against Data Augmentation
by: Gong, Xueluan, et al.
Published: (2025)
by: Gong, Xueluan, et al.
Published: (2025)
Parameter Training Efficiency Aware Resource Allocation for AIGC in Space-Air-Ground Integrated Networks
by: Qian, Liangxin, et al.
Published: (2024)
by: Qian, Liangxin, et al.
Published: (2024)
Overcoming Intrinsic Dispersion Locking for Achieving Spatio-Spectral Selectivity with Misaligned Bi-metagratings
by: Zhuang, Ze-Peng, et al.
Published: (2025)
by: Zhuang, Ze-Peng, et al.
Published: (2025)
Empowering Generalist Material Intelligence with Large Language Models
by: Wenhao Yuan, et al.
Published: (2025)
by: Wenhao Yuan, et al.
Published: (2025)
Forewarned is Forearmed: A Survey on Large Language Model-based Agents in Autonomous Cyberattacks
by: Xu, Minrui, et al.
Published: (2025)
by: Xu, Minrui, et al.
Published: (2025)
A Learning-based Incentive Mechanism for Mobile AIGC Service in Decentralized Internet of Vehicles
by: Fan, Jiani, et al.
Published: (2024)
by: Fan, Jiani, et al.
Published: (2024)
Similar Items
-
Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution
by: Chen, Chen, et al.
Published: (2025) -
AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions
by: Chen, Chen, et al.
Published: (2024) -
Neutralizing Backdoors through Information Conflicts for Large Language Models
by: Chen, Chen, et al.
Published: (2024) -
Threats, Attacks, and Defenses in Machine Unlearning: A Survey
by: Liu, Ziyao, et al.
Published: (2024) -
Attribution Techniques for Mitigating Hallucinated Information in RAG Systems: A Survey
by: Zhao, Yuqing, et al.
Published: (2026)