Democratizing Reward Design for Personal and Representative Value-Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Blair, Carter, Larson, Kate, Law, Edith
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914997004664832
author Blair, Carter
Larson, Kate
Law, Edith
author_facet Blair, Carter
Larson, Kate
Law, Edith
contents Aligning AI agents with human values is challenging due to diverse and subjective notions of values. Standard alignment methods often aggregate crowd feedback, which can result in the suppression of unique or minority preferences. We introduce Interactive-Reflective Dialogue Alignment, a method that iteratively engages users in reflecting on and specifying their subjective value definitions. This system learns individual value definitions through language-model-based preference elicitation and constructs personalized reward models that can be used to align AI behaviour. We evaluated our system through two studies with 30 participants, one focusing on "respect" and the other on ethical decision-making in autonomous vehicles. Our findings demonstrate diverse definitions of value-aligned behaviour and show that our system can accurately capture each person's unique understanding. This approach enables personalized alignment and can inform more representative and interpretable collective alignment strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2410_22203
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Democratizing Reward Design for Personal and Representative Value-Alignment
Blair, Carter
Larson, Kate
Law, Edith
Artificial Intelligence
Human-Computer Interaction
Aligning AI agents with human values is challenging due to diverse and subjective notions of values. Standard alignment methods often aggregate crowd feedback, which can result in the suppression of unique or minority preferences. We introduce Interactive-Reflective Dialogue Alignment, a method that iteratively engages users in reflecting on and specifying their subjective value definitions. This system learns individual value definitions through language-model-based preference elicitation and constructs personalized reward models that can be used to align AI behaviour. We evaluated our system through two studies with 30 participants, one focusing on "respect" and the other on ethical decision-making in autonomous vehicles. Our findings demonstrate diverse definitions of value-aligned behaviour and show that our system can accurately capture each person's unique understanding. This approach enables personalized alignment and can inform more representative and interpretable collective alignment strategies.
title Democratizing Reward Design for Personal and Representative Value-Alignment
topic Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2410.22203