The Coming Crisis of Multi-Agent Misalignment: AI Alignment Must Be a Dynamic and Social Process
Fuente:
arXiv
Saved in:
| Main Authors: | Carichon, Florian, Khandelwal, Aditi, Fauchard, Marylou, Farnadi, Golnoosh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reasoning with Preference Constraints: A Benchmark for Language Models in Many-to-One Matching Markets
by: Fauchard, Marylou, et al.
Published: (2025)
by: Fauchard, Marylou, et al.
Published: (2025)
Embedding Cultural Diversity in Prototype-based Recommender Systems
by: Moradi, Armin, et al.
Published: (2024)
by: Moradi, Armin, et al.
Published: (2024)
Understanding Intrinsic Socioeconomic Biases in Large Language Models
by: Arzaghi, Mina, et al.
Published: (2024)
by: Arzaghi, Mina, et al.
Published: (2024)
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
by: Farashah, Alireza Dehghanpour, et al.
Published: (2026)
by: Farashah, Alireza Dehghanpour, et al.
Published: (2026)
Intrinsic Meets Extrinsic Fairness: Assessing the Downstream Impact of Bias Mitigation in Large Language Models
by: Arzaghi', 'Mina, et al.
Published: (2025)
by: Arzaghi', 'Mina, et al.
Published: (2025)
Fairness in Federated Learning: Fairness for Whom?
by: Taik, Afaf, et al.
Published: (2025)
by: Taik, Afaf, et al.
Published: (2025)
Position: Cracking the Code of Cascading Disparity Towards Marginalized Communities
by: Farnadi, Golnoosh, et al.
Published: (2024)
by: Farnadi, Golnoosh, et al.
Published: (2024)
Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs
by: Pearman, Edie, et al.
Published: (2026)
by: Pearman, Edie, et al.
Published: (2026)
Different Horses for Different Courses: Comparing Bias Mitigation Algorithms in ML
by: Ganesh, Prakhar, et al.
Published: (2024)
by: Ganesh, Prakhar, et al.
Published: (2024)
Crossing Boundaries: Leveraging Semantic Divergences to Explore Cultural Novelty in Cooking Recipes
by: Carichon, Florian, et al.
Published: (2025)
by: Carichon, Florian, et al.
Published: (2025)
The Cost of Arbitrariness for Individuals: Examining the Legal and Technical Challenges of Model Multiplicity
by: Ganesh, Prakhar, et al.
Published: (2024)
by: Ganesh, Prakhar, et al.
Published: (2024)
LoRA Provides Differential Privacy by Design via Random Sketching
by: Malekmohammadi, Saber, et al.
Published: (2024)
by: Malekmohammadi, Saber, et al.
Published: (2024)
Some[Body] Must Receive That Pain for Agent Accountability
by: Hu, Botao Amber, et al.
Published: (2026)
by: Hu, Botao Amber, et al.
Published: (2026)
To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack
by: Zhuo, Terry Yue, et al.
Published: (2026)
by: Zhuo, Terry Yue, et al.
Published: (2026)
Can LLMs Cook Jamaican Couscous? A Study of Cultural Novelty in Recipe Generation
by: Carichon, F., et al.
Published: (2026)
by: Carichon, F., et al.
Published: (2026)
Synthetic Data for Robust AI Model Development in Regulated Enterprises
by: Godbole, Aditi
Published: (2025)
by: Godbole, Aditi
Published: (2025)
Agentic AI in Healthcare & Medicine: A Seven-Dimensional Taxonomy for Empirical Evaluation of LLM-based Agents
by: Vatsal, Shubham, et al.
Published: (2026)
by: Vatsal, Shubham, et al.
Published: (2026)
Understanding the Process of Human-AI Value Alignment
by: McKinlay, Jack, et al.
Published: (2025)
by: McKinlay, Jack, et al.
Published: (2025)
From Sycophantic Consensus to Pluralistic Repair: Why AI Alignment Must Surface Disagreement
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
Misaligned by Reward: Socially Undesirable Preferences in LLMs
by: Ghazaryan, Gayane, et al.
Published: (2026)
by: Ghazaryan, Gayane, et al.
Published: (2026)
Human Attribution of Causality to AI Across Agency, Misuse, and Misalignment
by: Carro, Maria Victoria, et al.
Published: (2026)
by: Carro, Maria Victoria, et al.
Published: (2026)
Towards an AI Observatory for the Nuclear Sector: A tool for anticipatory governance
by: Verma, Aditi, et al.
Published: (2025)
by: Verma, Aditi, et al.
Published: (2025)
How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures
by: Zhang, Shan, et al.
Published: (2026)
by: Zhang, Shan, et al.
Published: (2026)
AI based Content Creation and Product Recommendation Applications in E-commerce: An Ethical overview
by: Jain, Aditi Madhusudan, et al.
Published: (2025)
by: Jain, Aditi Madhusudan, et al.
Published: (2025)
Characterizing AI Agents for Alignment and Governance
by: Kasirzadeh, Atoosa, et al.
Published: (2025)
by: Kasirzadeh, Atoosa, et al.
Published: (2025)
Explainable AI Systems Must Be Contestable: Here's How to Make It Happen
by: Moreira, Catarina, et al.
Published: (2025)
by: Moreira, Catarina, et al.
Published: (2025)
Anchoring AI Capabilities in Market Valuations: The Capability Realization Rate Model and Valuation Misalignment Risk
by: Fang, Xinmin, et al.
Published: (2025)
by: Fang, Xinmin, et al.
Published: (2025)
Towards Computational Social Dynamics of Semi-Autonomous AI Agents
by: Lidarity, S. O., et al.
Published: (2026)
by: Lidarity, S. O., et al.
Published: (2026)
Misalignment or misuse? The AGI alignment tradeoff
by: Hellrigel-Holderbaum, Max, et al.
Published: (2025)
by: Hellrigel-Holderbaum, Max, et al.
Published: (2025)
Self-Explanation in Social AI Agents
by: Basappa, Rhea, et al.
Published: (2025)
by: Basappa, Rhea, et al.
Published: (2025)
Conformity and Social Impact on AI Agents
by: Bellina, Alessandro, et al.
Published: (2026)
by: Bellina, Alessandro, et al.
Published: (2026)
Too Human to Model:The Uncanny Valley of LLMs in Social Simulation -- When Generative Language Agents Misalign with Modelling Principles
by: Zeng, Yongchao, et al.
Published: (2025)
by: Zeng, Yongchao, et al.
Published: (2025)
The AI Alignment Paradox
by: West, Robert, et al.
Published: (2024)
by: West, Robert, et al.
Published: (2024)
Social Catalysts, Not Moral Agents: The Illusion of Alignment in LLM Societies
by: Hu, Yueqing, et al.
Published: (2026)
by: Hu, Yueqing, et al.
Published: (2026)
Ten Hard Problems in Artificial Intelligence We Must Get Right
by: Leech, Gavin, et al.
Published: (2024)
by: Leech, Gavin, et al.
Published: (2024)
Must Read: A Comprehensive Survey of Computational Persuasion
by: Bozdag, Nimet Beyza, et al.
Published: (2025)
by: Bozdag, Nimet Beyza, et al.
Published: (2025)
From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents
by: Qu, Jinxian, et al.
Published: (2026)
by: Qu, Jinxian, et al.
Published: (2026)
Rethinking AI Cultural Alignment
by: Bravansky, Michal, et al.
Published: (2025)
by: Bravansky, Michal, et al.
Published: (2025)
Multi-Agent LLMs as Ethics Advocates for AI-Based Systems
by: Yamani, Asma, et al.
Published: (2025)
by: Yamani, Asma, et al.
Published: (2025)
DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow
by: Long, Tao, et al.
Published: (2025)
by: Long, Tao, et al.
Published: (2025)
Similar Items
-
Reasoning with Preference Constraints: A Benchmark for Language Models in Many-to-One Matching Markets
by: Fauchard, Marylou, et al.
Published: (2025) -
Embedding Cultural Diversity in Prototype-based Recommender Systems
by: Moradi, Armin, et al.
Published: (2024) -
Understanding Intrinsic Socioeconomic Biases in Large Language Models
by: Arzaghi, Mina, et al.
Published: (2024) -
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
by: Farashah, Alireza Dehghanpour, et al.
Published: (2026) -
Intrinsic Meets Extrinsic Fairness: Assessing the Downstream Impact of Bias Mitigation in Large Language Models
by: Arzaghi', 'Mina, et al.
Published: (2025)