Misalignment or misuse? The AGI alignment tradeoff
Fuente:
arXiv
Saved in:
| Main Authors: | Hellrigel-Holderbaum, Max, Dung, Leonard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Against racing to AGI: Cooperation, deterrence, and catastrophic risks
by: Dung, Leonard, et al.
Published: (2025)
by: Dung, Leonard, et al.
Published: (2025)
Questionnaire Responses Do not Capture the Safety of AI Agents
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors
by: Wiedermann-Möller, Jonas, et al.
Published: (2026)
by: Wiedermann-Möller, Jonas, et al.
Published: (2026)
From Checklists to Clusters: A Homeostatic Account of AGI Evaluation
by: Reynolds, Brett
Published: (2025)
by: Reynolds, Brett
Published: (2025)
An Approach to Technical AGI Safety and Security
by: Shah, Rohin, et al.
Published: (2025)
by: Shah, Rohin, et al.
Published: (2025)
Are we Doomed to an AI Race? Why Self-Interest Could Drive Countries Towards a Moratorium on Superintelligence
by: Roussel, Edward, et al.
Published: (2026)
by: Roussel, Edward, et al.
Published: (2026)
Human Attribution of Causality to AI Across Agency, Misuse, and Misalignment
by: Carro, Maria Victoria, et al.
Published: (2026)
by: Carro, Maria Victoria, et al.
Published: (2026)
Misaligned by Reward: Socially Undesirable Preferences in LLMs
by: Ghazaryan, Gayane, et al.
Published: (2026)
by: Ghazaryan, Gayane, et al.
Published: (2026)
Why AI Alignment Failure Is Structural: Learned Human Interaction Structures and AGI as an Endogenous Evolutionary Shock
by: Sornette, Didier, et al.
Published: (2026)
by: Sornette, Didier, et al.
Published: (2026)
How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures
by: Zhang, Shan, et al.
Published: (2026)
by: Zhang, Shan, et al.
Published: (2026)
The Coming Crisis of Multi-Agent Misalignment: AI Alignment Must Be a Dynamic and Social Process
by: Carichon, Florian, et al.
Published: (2025)
by: Carichon, Florian, et al.
Published: (2025)
Anchoring AI Capabilities in Market Valuations: The Capability Realization Rate Model and Valuation Misalignment Risk
by: Fang, Xinmin, et al.
Published: (2025)
by: Fang, Xinmin, et al.
Published: (2025)
"They parted illusions -- they parted disclaim marinade": Misalignment as structural fidelity in LLMs
by: Costa, Mariana Lins
Published: (2025)
by: Costa, Mariana Lins
Published: (2025)
How Far Are We From AGI: Are LLMs All We Need?
by: Feng, Tao, et al.
Published: (2024)
by: Feng, Tao, et al.
Published: (2024)
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
by: Yang, Chao, et al.
Published: (2024)
by: Yang, Chao, et al.
Published: (2024)
Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
by: Williams, Kai, et al.
Published: (2025)
by: Williams, Kai, et al.
Published: (2025)
Moloch's Bargain: Emergent Misalignment When LLMs Compete for Audiences
by: El, Batu, et al.
Published: (2025)
by: El, Batu, et al.
Published: (2025)
What do people expect from Artificial Intelligence? Public opinion on alignment in AI moderation from Germany and the United States
by: Jungherr, Andreas, et al.
Published: (2025)
by: Jungherr, Andreas, et al.
Published: (2025)
Algorithms for learning value-aligned policies considering admissibility relaxation
by: Holgado-Sánchez, Andrés, et al.
Published: (2024)
by: Holgado-Sánchez, Andrés, et al.
Published: (2024)
BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format
by: Pihlakas, Roland, et al.
Published: (2025)
by: Pihlakas, Roland, et al.
Published: (2025)
Too Human to Model:The Uncanny Valley of LLMs in Social Simulation -- When Generative Language Agents Misalign with Modelling Principles
by: Zeng, Yongchao, et al.
Published: (2025)
by: Zeng, Yongchao, et al.
Published: (2025)
Investigating social alignment via mirroring in a system of interacting language models
by: McGuinness, Harvey, et al.
Published: (2024)
by: McGuinness, Harvey, et al.
Published: (2024)
AI-University: An LLM-based platform for instructional alignment to scientific classrooms
by: Shojaei, Mostafa Faghih, et al.
Published: (2025)
by: Shojaei, Mostafa Faghih, et al.
Published: (2025)
A Multi-Level Strategy for Deepfake Content Moderation under EU Regulation
by: Förster, Max-Paul, et al.
Published: (2025)
by: Förster, Max-Paul, et al.
Published: (2025)
Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations
by: Shrivastava, Aryan, et al.
Published: (2024)
by: Shrivastava, Aryan, et al.
Published: (2024)
Risks from Language Models for Automated Mental Healthcare: Ethics and Structure for Implementation
by: Grabb, Declan, et al.
Published: (2024)
by: Grabb, Declan, et al.
Published: (2024)
Individual utilities of life satisfaction reveal inequality aversion unrelated to political alignment
by: Cooper, Crispin, et al.
Published: (2025)
by: Cooper, Crispin, et al.
Published: (2025)
Jolting Technologies: Superexponential Acceleration in AI Capabilities and Implications for AGI
by: Orban, David
Published: (2025)
by: Orban, David
Published: (2025)
Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models
by: Potham, Ram, et al.
Published: (2025)
by: Potham, Ram, et al.
Published: (2025)
Building Capacity for Artificial Intelligence in Africa: A Cross-Country Survey of Challenges and Governance Pathways
by: Aryee, Jeffrey N. A., et al.
Published: (2025)
by: Aryee, Jeffrey N. A., et al.
Published: (2025)
The Arrival of AGI? When Expert Personas Exceed Expert Benchmarks
by: Mullens, Drake, et al.
Published: (2026)
by: Mullens, Drake, et al.
Published: (2026)
When Is It Acceptable to Break the Rules? Knowledge Representation of Moral Judgement Based on Empirical Data
by: Awad, Edmond, et al.
Published: (2022)
by: Awad, Edmond, et al.
Published: (2022)
High vs. Low AGI: Ontology and Conceptual Taxonomy for Geopolitical Coherence
by: Max, Antonio
Published: (2025)
by: Max, Antonio
Published: (2025)
The bitter lesson of misuse detection
by: Mariaccia, Hadrien, et al.
Published: (2025)
by: Mariaccia, Hadrien, et al.
Published: (2025)
What are human values, and how do we align AI to them?
by: Klingefjord, Oliver, et al.
Published: (2024)
by: Klingefjord, Oliver, et al.
Published: (2024)
Human vs. Machine: Behavioral Differences Between Expert Humans and Language Models in Wargame Simulations
by: Lamparth, Max, et al.
Published: (2024)
by: Lamparth, Max, et al.
Published: (2024)
Visibility into AI Agents
by: Chan, Alan, et al.
Published: (2024)
by: Chan, Alan, et al.
Published: (2024)
Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms
by: Hu, Botao Amber, et al.
Published: (2026)
by: Hu, Botao Amber, et al.
Published: (2026)
Similar Items
-
Against racing to AGI: Cooperation, deterrence, and catastrophic risks
by: Dung, Leonard, et al.
Published: (2025) -
Questionnaire Responses Do not Capture the Safety of AI Agents
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026) -
Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors
by: Wiedermann-Möller, Jonas, et al.
Published: (2026) -
From Checklists to Clusters: A Homeostatic Account of AGI Evaluation
by: Reynolds, Brett
Published: (2025) -
An Approach to Technical AGI Safety and Security
by: Shah, Rohin, et al.
Published: (2025)