Culturally-Attuned Moral Machines: Implicit Learning of Human Value Systems by AI through Inverse Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Oliveira, Nigini, Li, Jasmine, Khalvati, Koosha, Barragan, Rodolfo Cortes, Reinecke, Katharina, Meltzoff, Andrew N., Rao, Rajesh P. N.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917170713198592
author Oliveira, Nigini
Li, Jasmine
Khalvati, Koosha
Barragan, Rodolfo Cortes
Reinecke, Katharina
Meltzoff, Andrew N.
Rao, Rajesh P. N.
author_facet Oliveira, Nigini
Li, Jasmine
Khalvati, Koosha
Barragan, Rodolfo Cortes
Reinecke, Katharina
Meltzoff, Andrew N.
Rao, Rajesh P. N.
contents Constructing a universal moral code for artificial intelligence (AI) is difficult or even impossible, given that different human cultures have different definitions of morality and different societal norms. We therefore argue that the value system of an AI should be culturally attuned: just as a child raised in a particular culture learns the specific values and norms of that culture, we propose that an AI agent operating in a particular human community should acquire that community's moral, ethical, and cultural codes. How AI systems might acquire such codes from human observation and interaction has remained an open question. Here, we propose using inverse reinforcement learning (IRL) as a method for AI agents to acquire a culturally-attuned value system implicitly. We test our approach using an experimental paradigm in which AI agents use IRL to learn different reward functions, which govern the agents' moral values, by observing the behavior of different cultural groups in an online virtual world requiring real-time decision making. We show that an AI agent learning from the average behavior of a particular cultural group can acquire altruistic characteristics reflective of that group's behavior, and this learned value system can generalize to new scenarios requiring altruistic judgments. Our results provide, to our knowledge, the first demonstration that AI agents could potentially be endowed with the ability to continually learn their values and norms from observing and interacting with humans, thereby becoming attuned to the culture they are operating in.
format Preprint
id arxiv_https___arxiv_org_abs_2312_17479
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Culturally-Attuned Moral Machines: Implicit Learning of Human Value Systems by AI through Inverse Reinforcement Learning
Oliveira, Nigini
Li, Jasmine
Khalvati, Koosha
Barragan, Rodolfo Cortes
Reinecke, Katharina
Meltzoff, Andrew N.
Rao, Rajesh P. N.
Artificial Intelligence
Computers and Society
Human-Computer Interaction
Machine Learning
Constructing a universal moral code for artificial intelligence (AI) is difficult or even impossible, given that different human cultures have different definitions of morality and different societal norms. We therefore argue that the value system of an AI should be culturally attuned: just as a child raised in a particular culture learns the specific values and norms of that culture, we propose that an AI agent operating in a particular human community should acquire that community's moral, ethical, and cultural codes. How AI systems might acquire such codes from human observation and interaction has remained an open question. Here, we propose using inverse reinforcement learning (IRL) as a method for AI agents to acquire a culturally-attuned value system implicitly. We test our approach using an experimental paradigm in which AI agents use IRL to learn different reward functions, which govern the agents' moral values, by observing the behavior of different cultural groups in an online virtual world requiring real-time decision making. We show that an AI agent learning from the average behavior of a particular cultural group can acquire altruistic characteristics reflective of that group's behavior, and this learned value system can generalize to new scenarios requiring altruistic judgments. Our results provide, to our knowledge, the first demonstration that AI agents could potentially be endowed with the ability to continually learn their values and norms from observing and interacting with humans, thereby becoming attuned to the culture they are operating in.
title Culturally-Attuned Moral Machines: Implicit Learning of Human Value Systems by AI through Inverse Reinforcement Learning
topic Artificial Intelligence
Computers and Society
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2312.17479