Against racing to AGI: Cooperation, deterrence, and catastrophic risks
Fuente:
arXiv
Saved in:
| Main Authors: | Dung, Leonard, Hellrigel-Holderbaum, Max |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Misalignment or misuse? The AGI alignment tradeoff
by: Hellrigel-Holderbaum, Max, et al.
Published: (2025)
by: Hellrigel-Holderbaum, Max, et al.
Published: (2025)
Questionnaire Responses Do not Capture the Safety of AI Agents
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
What AI evaluations for preventing catastrophic risks can and cannot do
by: Barnett, Peter, et al.
Published: (2024)
by: Barnett, Peter, et al.
Published: (2024)
Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors
by: Wiedermann-Möller, Jonas, et al.
Published: (2026)
by: Wiedermann-Möller, Jonas, et al.
Published: (2026)
From Checklists to Clusters: A Homeostatic Account of AGI Evaluation
by: Reynolds, Brett
Published: (2025)
by: Reynolds, Brett
Published: (2025)
An Approach to Technical AGI Safety and Security
by: Shah, Rohin, et al.
Published: (2025)
by: Shah, Rohin, et al.
Published: (2025)
Are we Doomed to an AI Race? Why Self-Interest Could Drive Countries Towards a Moratorium on Superintelligence
by: Roussel, Edward, et al.
Published: (2026)
by: Roussel, Edward, et al.
Published: (2026)
Why AI Alignment Failure Is Structural: Learned Human Interaction Structures and AGI as an Endogenous Evolutionary Shock
by: Sornette, Didier, et al.
Published: (2026)
by: Sornette, Didier, et al.
Published: (2026)
How Far Are We From AGI: Are LLMs All We Need?
by: Feng, Tao, et al.
Published: (2024)
by: Feng, Tao, et al.
Published: (2024)
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
by: Yang, Chao, et al.
Published: (2024)
by: Yang, Chao, et al.
Published: (2024)
Defending Compute Thresholds Against Legal Loopholes
by: Pistillo, Matteo, et al.
Published: (2025)
by: Pistillo, Matteo, et al.
Published: (2025)
Affirmative safety: An approach to risk management for high-risk AI
by: Wasil, Akash R., et al.
Published: (2024)
by: Wasil, Akash R., et al.
Published: (2024)
Contemporary AI foundation models increase biological weapons risk
by: Brent, Roger, et al.
Published: (2025)
by: Brent, Roger, et al.
Published: (2025)
US-China perspectives on extreme AI risks and global governance
by: Wasil, Akash, et al.
Published: (2024)
by: Wasil, Akash, et al.
Published: (2024)
Episodic memory in AI agents poses risks that should be studied and mitigated
by: DeChant, Chad
Published: (2025)
by: DeChant, Chad
Published: (2025)
Human services organizations and the responsible integration of AI: Considering ethics and contextualizing risk(s)
by: Perron, Brian E., et al.
Published: (2025)
by: Perron, Brian E., et al.
Published: (2025)
A Multi-Level Strategy for Deepfake Content Moderation under EU Regulation
by: Förster, Max-Paul, et al.
Published: (2025)
by: Förster, Max-Paul, et al.
Published: (2025)
"Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness
by: Karr Jr., Jonathan A., et al.
Published: (2025)
by: Karr Jr., Jonathan A., et al.
Published: (2025)
POEX: Towards Policy Executable Jailbreak Attacks Against the LLM-based Robots
by: Lu, Xuancun, et al.
Published: (2024)
by: Lu, Xuancun, et al.
Published: (2024)
Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations
by: Shrivastava, Aryan, et al.
Published: (2024)
by: Shrivastava, Aryan, et al.
Published: (2024)
Risks from Language Models for Automated Mental Healthcare: Ethics and Structure for Implementation
by: Grabb, Declan, et al.
Published: (2024)
by: Grabb, Declan, et al.
Published: (2024)
Conversational Agents for Building Energy Efficiency -- Advising Housing Cooperatives in Stockholm on Reducing Energy Consumption
by: Ghani, Shadaab, et al.
Published: (2025)
by: Ghani, Shadaab, et al.
Published: (2025)
Jolting Technologies: Superexponential Acceleration in AI Capabilities and Implications for AGI
by: Orban, David
Published: (2025)
by: Orban, David
Published: (2025)
Governing frontier general-purpose AI in the public sector: adaptive risk management and policy capacity under uncertainty through 2030
by: Xavier, Fabio Correa
Published: (2026)
by: Xavier, Fabio Correa
Published: (2026)
Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models
by: Potham, Ram, et al.
Published: (2025)
by: Potham, Ram, et al.
Published: (2025)
Evaluating Large Language Models Against Human Annotators in Latent Content Analysis: Sentiment, Political Leaning, Emotional Intensity, and Sarcasm
by: Bojic, Ljubisa, et al.
Published: (2025)
by: Bojic, Ljubisa, et al.
Published: (2025)
The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analyses of AI safety policies
by: Coggins, Sam, et al.
Published: (2025)
by: Coggins, Sam, et al.
Published: (2025)
Creating a Cooperative AI Policymaking Platform through Open Source Collaboration
by: Lewington, Aiden, et al.
Published: (2024)
by: Lewington, Aiden, et al.
Published: (2024)
To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack
by: Zhuo, Terry Yue, et al.
Published: (2026)
by: Zhuo, Terry Yue, et al.
Published: (2026)
Building Capacity for Artificial Intelligence in Africa: A Cross-Country Survey of Challenges and Governance Pathways
by: Aryee, Jeffrey N. A., et al.
Published: (2025)
by: Aryee, Jeffrey N. A., et al.
Published: (2025)
Embodied LLM Agents Learn to Cooperate in Organized Teams
by: Guo, Xudong, et al.
Published: (2024)
by: Guo, Xudong, et al.
Published: (2024)
The Arrival of AGI? When Expert Personas Exceed Expert Benchmarks
by: Mullens, Drake, et al.
Published: (2026)
by: Mullens, Drake, et al.
Published: (2026)
When Is It Acceptable to Break the Rules? Knowledge Representation of Moral Judgement Based on Empirical Data
by: Awad, Edmond, et al.
Published: (2022)
by: Awad, Edmond, et al.
Published: (2022)
The global consensus on the risk management of autonomous driving
by: Krügel, Sebastian, et al.
Published: (2025)
by: Krügel, Sebastian, et al.
Published: (2025)
High vs. Low AGI: Ontology and Conceptual Taxonomy for Geopolitical Coherence
by: Max, Antonio
Published: (2025)
by: Max, Antonio
Published: (2025)
Token Taxes: mitigating AGI's economic risks
by: Irwin, Lucas, et al.
Published: (2026)
by: Irwin, Lucas, et al.
Published: (2026)
Human vs. Machine: Behavioral Differences Between Expert Humans and Language Models in Wargame Simulations
by: Lamparth, Max, et al.
Published: (2024)
by: Lamparth, Max, et al.
Published: (2024)
Visibility into AI Agents
by: Chan, Alan, et al.
Published: (2024)
by: Chan, Alan, et al.
Published: (2024)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
by: Xiao, Yuxin, et al.
Published: (2025)
by: Xiao, Yuxin, et al.
Published: (2025)
Artificial intelligence and the Gulf Cooperation Council workforce adapting to the future of work
by: Albous, Mohammad Rashed, et al.
Published: (2025)
by: Albous, Mohammad Rashed, et al.
Published: (2025)
Similar Items
-
Misalignment or misuse? The AGI alignment tradeoff
by: Hellrigel-Holderbaum, Max, et al.
Published: (2025) -
Questionnaire Responses Do not Capture the Safety of AI Agents
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026) -
What AI evaluations for preventing catastrophic risks can and cannot do
by: Barnett, Peter, et al.
Published: (2024) -
Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors
by: Wiedermann-Möller, Jonas, et al.
Published: (2026) -
From Checklists to Clusters: A Homeostatic Account of AGI Evaluation
by: Reynolds, Brett
Published: (2025)