Learning Human-like Representations to Enable Learning Human Values

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wynn, Andrea, Sucholutsky, Ilia, Griffiths, Thomas L.
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908124425748480
author Wynn, Andrea
Sucholutsky, Ilia
Griffiths, Thomas L.
author_facet Wynn, Andrea
Sucholutsky, Ilia
Griffiths, Thomas L.
contents How can we build AI systems that can learn any set of individual human values both quickly and safely, avoiding causing harm or violating societal standards for acceptable behavior during the learning process? We explore the effects of representational alignment between humans and AI agents on learning human values. Making AI systems learn human-like representations of the world has many known benefits, including improving generalization, robustness to domain shifts, and few-shot learning performance. We demonstrate that this kind of representational alignment can also support safely learning and exploring human values in the context of personalization. We begin with a theoretical prediction, show that it applies to learning human morality judgments, then show that our results generalize to ten different aspects of human values -- including ethics, honesty, and fairness -- training AI agents on each set of values in a multi-armed bandit setting, where rewards reflect human value judgments over the chosen action. Using a set of textual action descriptions, we collect value judgments from humans, as well as similarity judgments from both humans and multiple language models, and demonstrate that representational alignment enables both safe exploration and improved generalization when learning human values.
format Preprint
id arxiv_https___arxiv_org_abs_2312_14106
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learning Human-like Representations to Enable Learning Human Values
Wynn, Andrea
Sucholutsky, Ilia
Griffiths, Thomas L.
Artificial Intelligence
Machine Learning
How can we build AI systems that can learn any set of individual human values both quickly and safely, avoiding causing harm or violating societal standards for acceptable behavior during the learning process? We explore the effects of representational alignment between humans and AI agents on learning human values. Making AI systems learn human-like representations of the world has many known benefits, including improving generalization, robustness to domain shifts, and few-shot learning performance. We demonstrate that this kind of representational alignment can also support safely learning and exploring human values in the context of personalization. We begin with a theoretical prediction, show that it applies to learning human morality judgments, then show that our results generalize to ten different aspects of human values -- including ethics, honesty, and fairness -- training AI agents on each set of values in a multi-armed bandit setting, where rewards reflect human value judgments over the chosen action. Using a set of textual action descriptions, we collect value judgments from humans, as well as similarity judgments from both humans and multiple language models, and demonstrate that representational alignment enables both safe exploration and improved generalization when learning human values.
title Learning Human-like Representations to Enable Learning Human Values
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2312.14106