Persona Features Control Emergent Misalignment
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Miles, la Tour, Tom Dupré, Watkins, Olivia, Makelov, Alex, Chi, Ryan A., Miserendino, Samuel, Wang, Jeffrey, Rajaram, Achyuta, Heidecke, Johannes, Patwardhan, Tejal, Mossing, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking
by: Ustaomeroglu, Muhammed, et al.
Published: (2026)
by: Ustaomeroglu, Muhammed, et al.
Published: (2026)
BudgetMLAgent: A Cost-Effective LLM Multi-Agent system for Automating Machine Learning Tasks
by: Gandhi, Shubham, et al.
Published: (2024)
by: Gandhi, Shubham, et al.
Published: (2024)
Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English
by: Anderson, Bryce, et al.
Published: (2025)
by: Anderson, Bryce, et al.
Published: (2025)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
by: Wang, Zhen, et al.
Published: (2025)
by: Wang, Zhen, et al.
Published: (2025)
One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents
by: Hong, Yoosung
Published: (2026)
by: Hong, Yoosung
Published: (2026)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
by: Du, Bangde, et al.
Published: (2025)
by: Du, Bangde, et al.
Published: (2025)
Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors
by: Williamson, Dane, et al.
Published: (2025)
by: Williamson, Dane, et al.
Published: (2025)
The Arrival of AGI? When Expert Personas Exceed Expert Benchmarks
by: Mullens, Drake, et al.
Published: (2026)
by: Mullens, Drake, et al.
Published: (2026)
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
by: Naik, Akshat, et al.
Published: (2025)
by: Naik, Akshat, et al.
Published: (2025)
FACT: Multinomial Misalignment Classification for Point Cloud Registration
by: Dillén, Ludvig, et al.
Published: (2025)
by: Dillén, Ludvig, et al.
Published: (2025)
The Pragmatic Persona: Discovering LLM Persona through Bridging Inference
by: Yang, Jisoo, et al.
Published: (2026)
by: Yang, Jisoo, et al.
Published: (2026)
Aligning by Misaligning: Boundary-aware Curriculum Learning for Multimodal Alignment
by: Ye, Hua, et al.
Published: (2025)
by: Ye, Hua, et al.
Published: (2025)
HyperPersona: A Multi-Level Hypergraph Framework for Text-Based Automatic Personality Prediction
by: Heydari, Sina, et al.
Published: (2026)
by: Heydari, Sina, et al.
Published: (2026)
OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models
by: Koroglu, Mathis, et al.
Published: (2024)
by: Koroglu, Mathis, et al.
Published: (2024)
Learning the meanings of function words from grounded language using a visual question answering model
by: Portelance, Eva, et al.
Published: (2023)
by: Portelance, Eva, et al.
Published: (2023)
Trapped in texture bias? A large scale comparison of deep instance segmentation
by: Theodoridis, Johannes, et al.
Published: (2024)
by: Theodoridis, Johannes, et al.
Published: (2024)
Emergent Coordination in Multi-Agent Language Models
by: Riedl, Christoph
Published: (2025)
by: Riedl, Christoph
Published: (2025)
N-Agent Ad Hoc Teamwork
by: Wang, Caroline, et al.
Published: (2024)
by: Wang, Caroline, et al.
Published: (2024)
MESS+: Dynamically Learned Inference-Time LLM Routing in Model Zoos with Service Level Guarantees
by: Woisetschläger, Herbert, et al.
Published: (2025)
by: Woisetschläger, Herbert, et al.
Published: (2025)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
by: Kugler, Kai
Published: (2025)
by: Kugler, Kai
Published: (2025)
Motion Perceiver: Real-Time Occupancy Forecasting for Embedded Systems
by: Ferenczi, Bryce, et al.
Published: (2023)
by: Ferenczi, Bryce, et al.
Published: (2023)
From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents
by: Wang, Xinyue, et al.
Published: (2026)
by: Wang, Xinyue, et al.
Published: (2026)
Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback
by: Juzek, Tom S., et al.
Published: (2025)
by: Juzek, Tom S., et al.
Published: (2025)
What Would GPT Click: Practical Effects of Human-AI Behavioral Misalignment and the Cost of Synthetic Participants in User Experience
by: Kuric, Eduard, et al.
Published: (2026)
by: Kuric, Eduard, et al.
Published: (2026)
Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents
by: Chua, Jaymari, et al.
Published: (2025)
by: Chua, Jaymari, et al.
Published: (2025)
Advancing Multimodal Agent Reasoning with Long-Term Neuro-Symbolic Memory
by: Jiang, Rongjie, et al.
Published: (2026)
by: Jiang, Rongjie, et al.
Published: (2026)
FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation
by: Hildebrand, Samuel, et al.
Published: (2025)
by: Hildebrand, Samuel, et al.
Published: (2025)
SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
by: Miserendino, Samuel, et al.
Published: (2025)
by: Miserendino, Samuel, et al.
Published: (2025)
From Guessing to Asking: An Approach to Resolving the Persona Knowledge Gap in LLMs during Multi-Turn Conversations
by: Baskar, Sarvesh, et al.
Published: (2025)
by: Baskar, Sarvesh, et al.
Published: (2025)
Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustainable logging and post-harvest silvicultural treatment in tropical forest using airborne robotics systems
by: Albiero, Daniel, et al.
Published: (2026)
by: Albiero, Daniel, et al.
Published: (2026)
VeriPlan: Integrating Formal Verification and LLMs into End-User Planning
by: Lee, Christine, et al.
Published: (2025)
by: Lee, Christine, et al.
Published: (2025)
The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models
by: Wang, Zhixiang
Published: (2025)
by: Wang, Zhixiang
Published: (2025)
An Explainable Collaborative Dialogue System using a Theory of Mind
by: Cohen, Philip R., et al.
Published: (2023)
by: Cohen, Philip R., et al.
Published: (2023)
ROTATE: Regret-driven Open-ended Training for Ad Hoc Teamwork
by: Wang, Caroline, et al.
Published: (2025)
by: Wang, Caroline, et al.
Published: (2025)
Leveraging Large Language Models to Extract and Translate Medical Information in Doctors' Notes for Health Records and Diagnostic Billing Codes
by: Hartnett, Peter, et al.
Published: (2026)
by: Hartnett, Peter, et al.
Published: (2026)
Incentives for Responsiveness, Instrumental Control and Impact
by: Carey, Ryan, et al.
Published: (2020)
by: Carey, Ryan, et al.
Published: (2020)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent Memory
by: Gu, Yingjie, et al.
Published: (2026)
by: Gu, Yingjie, et al.
Published: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025)
by: Raoufi, Behnam, et al.
Published: (2025)
Similar Items
-
BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking
by: Ustaomeroglu, Muhammed, et al.
Published: (2026) -
BudgetMLAgent: A Cost-Effective LLM Multi-Agent system for Automating Machine Learning Tasks
by: Gandhi, Shubham, et al.
Published: (2024) -
Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English
by: Anderson, Bryce, et al.
Published: (2025) -
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
by: Wang, Zhen, et al.
Published: (2025) -
One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents
by: Hong, Yoosung
Published: (2026)