The Ultimate Test of Superintelligent AI Agents: Can an AI Balance Care and Control in Asymmetric Relationships?
Fuente:
arXiv
Saved in:
| Main Authors: | Bouneffouf, Djallel, Riemer, Matthew, Varshney, Kush |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scopes of Alignment
by: Varshney, Kush R., et al.
Published: (2025)
by: Varshney, Kush R., et al.
Published: (2025)
Assessing AI Utility: The Random Guesser Test for Sequential Decision-Making Systems
by: Ide, Shun, et al.
Published: (2024)
by: Ide, Shun, et al.
Published: (2024)
The Effectiveness of Approximate Regularized Replay for Efficient Supervised Fine-Tuning of Large Language Models
by: Riemer, Matthew, et al.
Published: (2025)
by: Riemer, Matthew, et al.
Published: (2025)
Survey: Multi-Armed Bandits Meet Large Language Models
by: Bouneffouf, Djallel, et al.
Published: (2025)
by: Bouneffouf, Djallel, et al.
Published: (2025)
Agentic AI Needs a Systems Theory
by: Miehling, Erik, et al.
Published: (2025)
by: Miehling, Erik, et al.
Published: (2025)
Contextual Moral Value Alignment Through Context-Based Aggregation
by: Dognin, Pierre, et al.
Published: (2024)
by: Dognin, Pierre, et al.
Published: (2024)
Position: Theory of Mind Benchmarks are Broken for Large Language Models
by: Riemer, Matthew, et al.
Published: (2024)
by: Riemer, Matthew, et al.
Published: (2024)
Decolonial AI Alignment: Openness, Viśe\d{s}a-Dharma, and Including Excluded Knowledges
by: Varshney, Kush R.
Published: (2023)
by: Varshney, Kush R.
Published: (2023)
Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?
by: Bengio, Yoshua, et al.
Published: (2025)
by: Bengio, Yoshua, et al.
Published: (2025)
An Algebraic Exposition of the Theory of Dyadic Morality
by: Varshney, Kush R.
Published: (2026)
by: Varshney, Kush R.
Published: (2026)
AI & Human Co-Improvement for Safer Co-Superintelligence
by: Weston, Jason, et al.
Published: (2025)
by: Weston, Jason, et al.
Published: (2025)
Mitigating Misalignment Contagion by Steering with Implicit Traits
by: Chang, Maria, et al.
Published: (2026)
by: Chang, Maria, et al.
Published: (2026)
Artificial Superintelligence May be Useless: Equilibria in the Economy of Multiple AI Agents
by: Cai, Huan, et al.
Published: (2026)
by: Cai, Huan, et al.
Published: (2026)
An Operational Kardashev-Style Scale for Autonomous AI - Towards AGI and Superintelligence
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
by: Kidder, William, et al.
Published: (2024)
by: Kidder, William, et al.
Published: (2024)
AI Steerability 360: A Toolkit for Steering Large Language Models
by: Miehling, Erik, et al.
Published: (2026)
by: Miehling, Erik, et al.
Published: (2026)
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
by: Huang, Zhen, et al.
Published: (2024)
by: Huang, Zhen, et al.
Published: (2024)
From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
The Einstein Test: Towards a Practical Test of a Machine's Ability to Exhibit Superintelligence
by: Benrimoh, David, et al.
Published: (2025)
by: Benrimoh, David, et al.
Published: (2025)
COMPASS: Computational Mapping of Patient-Therapist Alliance Strategies with Language Modeling
by: Lin, Baihan, et al.
Published: (2024)
by: Lin, Baihan, et al.
Published: (2024)
Computational Dualism and Objective Superintelligence
by: Bennett, Michael Timothy
Published: (2023)
by: Bennett, Michael Timothy
Published: (2023)
Superintelligent Retrieval Agent: The Next Frontier of Information Retrieval
by: Yang, Zeyu, et al.
Published: (2026)
by: Yang, Zeyu, et al.
Published: (2026)
Tool Building as a Path to "Superintelligence"
by: Koplow, David, et al.
Published: (2026)
by: Koplow, David, et al.
Published: (2026)
Superintelligence Strategy: Expert Version
by: Hendrycks, Dan, et al.
Published: (2025)
by: Hendrycks, Dan, et al.
Published: (2025)
Deconstructing Superintelligence: Identity, Self-Modification and Différance
by: Perrier, Elija
Published: (2026)
by: Perrier, Elija
Published: (2026)
Are we Doomed to an AI Race? Why Self-Interest Could Drive Countries Towards a Moratorium on Superintelligence
by: Roussel, Edward, et al.
Published: (2026)
by: Roussel, Edward, et al.
Published: (2026)
Hey GPT, Can You be More Racist? Analysis from Crowdsourced Attempts to Elicit Biased Content from Generative AI
by: Guo, Hangzhi, et al.
Published: (2024)
by: Guo, Hangzhi, et al.
Published: (2024)
Can We Test Consciousness Theories on AI? Ablations, Markers, and Robustness
by: Phua, Yin Jun
Published: (2025)
by: Phua, Yin Jun
Published: (2025)
Evaluating the Prompt Steerability of Large Language Models
by: Miehling, Erik, et al.
Published: (2024)
by: Miehling, Erik, et al.
Published: (2024)
Agent-Testing Agent: A Meta-Agent for Automated Testing and Evaluation of Conversational AI Agents
by: Komoravolu, Sameer, et al.
Published: (2025)
by: Komoravolu, Sameer, et al.
Published: (2025)
Can AI Master Econometrics? Evidence from Econometrics AI Agent on Expert-Level Tasks
by: Chen, Qiang, et al.
Published: (2025)
by: Chen, Qiang, et al.
Published: (2025)
Consolidation via Policy Information Regularization in Deep RL for Multi-Agent Games
by: Malloy, Tailia, et al.
Published: (2020)
by: Malloy, Tailia, et al.
Published: (2020)
Evaluating Control Protocols for Untrusted AI Agents
by: Kutasov, Jon, et al.
Published: (2025)
by: Kutasov, Jon, et al.
Published: (2025)
Containment Verification: AI Safety Guarantees Independent of Alignment
by: Moon, Royce, et al.
Published: (2026)
by: Moon, Royce, et al.
Published: (2026)
Multi-Agent Penetration Testing AI for the Web
by: David, Isaac, et al.
Published: (2025)
by: David, Isaac, et al.
Published: (2025)
From Predictions to Explanations: Explainable AI for Autism Diagnosis and Identification of Critical Brain Regions
by: Gupta, Kush, et al.
Published: (2025)
by: Gupta, Kush, et al.
Published: (2025)
Value-Sensitive AI for Prayer: Balancing the Agencies Between Human and AI Agents in Spiritual Context
by: Kwon, Soonho, et al.
Published: (2026)
by: Kwon, Soonho, et al.
Published: (2026)
The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
by: Overman, William, et al.
Published: (2025)
by: Overman, William, et al.
Published: (2025)
Can AI Agents Design and Implement Drug Discovery Pipelines?
by: Smbatyan, Khachik, et al.
Published: (2025)
by: Smbatyan, Khachik, et al.
Published: (2025)
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
Similar Items
-
Scopes of Alignment
by: Varshney, Kush R., et al.
Published: (2025) -
Assessing AI Utility: The Random Guesser Test for Sequential Decision-Making Systems
by: Ide, Shun, et al.
Published: (2024) -
The Effectiveness of Approximate Regularized Replay for Efficient Supervised Fine-Tuning of Large Language Models
by: Riemer, Matthew, et al.
Published: (2025) -
Survey: Multi-Armed Bandits Meet Large Language Models
by: Bouneffouf, Djallel, et al.
Published: (2025) -
Agentic AI Needs a Systems Theory
by: Miehling, Erik, et al.
Published: (2025)