Autonomous Alignment with Human Value on Altruism through Considerate Self-imagination and Theory of Mind

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tong, Haibo, Lu, Enmeng, Sun, Yinqian, Han, Zhengqiang, Liu, Chao, Zhao, Feifei, Zeng, Yi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913638942507008
author Tong, Haibo
Lu, Enmeng
Sun, Yinqian
Han, Zhengqiang
Liu, Chao
Zhao, Feifei
Zeng, Yi
author_facet Tong, Haibo
Lu, Enmeng
Sun, Yinqian
Han, Zhengqiang
Liu, Chao
Zhao, Feifei
Zeng, Yi
contents With the widespread application of Artificial Intelligence (AI) in human society, enabling AI to autonomously align with human values has become a pressing issue to ensure its sustainable development and benefit to humanity. One of the most important aspects of aligning with human values is the necessity for agents to autonomously make altruistic, safe, and ethical decisions, considering and caring for human well-being. Current AI extremely pursues absolute superiority in certain tasks, remaining indifferent to the surrounding environment and other agents, which has led to numerous safety risks. Altruistic behavior in human society originates from humans' capacity for empathizing others, known as Theory of Mind (ToM), combined with predictive imaginative interactions before taking action to produce thoughtful and altruistic behaviors. Inspired by this, we are committed to endow agents with considerate self-imagination and ToM capabilities, driving them through implicit intrinsic motivations to autonomously align with human altruistic values. By integrating ToM within the imaginative space, agents keep an eye on the well-being of other agents in real time, proactively anticipate potential risks to themselves and others, and make thoughtful altruistic decisions that balance negative effects on the environment. The ancient Chinese story of Sima Guang Smashes the Vat illustrates the moral behavior of the young Sima Guang smashed a vat to save a child who had accidentally fallen into it, which is an excellent reference scenario for this paper. We design an experimental scenario similar to Sima Guang Smashes the Vat and its variants with different complexities, which reflects the trade-offs and comprehensive considerations between self-goals, altruistic rescue, and avoiding negative side effects.
format Preprint
id arxiv_https___arxiv_org_abs_2501_00320
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Autonomous Alignment with Human Value on Altruism through Considerate Self-imagination and Theory of Mind
Tong, Haibo
Lu, Enmeng
Sun, Yinqian
Han, Zhengqiang
Liu, Chao
Zhao, Feifei
Zeng, Yi
Artificial Intelligence
With the widespread application of Artificial Intelligence (AI) in human society, enabling AI to autonomously align with human values has become a pressing issue to ensure its sustainable development and benefit to humanity. One of the most important aspects of aligning with human values is the necessity for agents to autonomously make altruistic, safe, and ethical decisions, considering and caring for human well-being. Current AI extremely pursues absolute superiority in certain tasks, remaining indifferent to the surrounding environment and other agents, which has led to numerous safety risks. Altruistic behavior in human society originates from humans' capacity for empathizing others, known as Theory of Mind (ToM), combined with predictive imaginative interactions before taking action to produce thoughtful and altruistic behaviors. Inspired by this, we are committed to endow agents with considerate self-imagination and ToM capabilities, driving them through implicit intrinsic motivations to autonomously align with human altruistic values. By integrating ToM within the imaginative space, agents keep an eye on the well-being of other agents in real time, proactively anticipate potential risks to themselves and others, and make thoughtful altruistic decisions that balance negative effects on the environment. The ancient Chinese story of Sima Guang Smashes the Vat illustrates the moral behavior of the young Sima Guang smashed a vat to save a child who had accidentally fallen into it, which is an excellent reference scenario for this paper. We design an experimental scenario similar to Sima Guang Smashes the Vat and its variants with different complexities, which reflects the trade-offs and comprehensive considerations between self-goals, altruistic rescue, and avoiding negative side effects.
title Autonomous Alignment with Human Value on Altruism through Considerate Self-imagination and Theory of Mind
topic Artificial Intelligence
url https://arxiv.org/abs/2501.00320