Mapping Social Choice Theory to RLHF

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Jessica, Fleisig, Eve
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909175876943872
author Dai, Jessica
Fleisig, Eve
author_facet Dai, Jessica
Fleisig, Eve
contents Recent work on the limitations of using reinforcement learning from human feedback (RLHF) to incorporate human preferences into model behavior often raises social choice theory as a reference point. Social choice theory's analysis of settings such as voting mechanisms provides technical infrastructure that can inform how to aggregate human preferences amid disagreement. We analyze the problem settings of social choice and RLHF, identify key differences between them, and discuss how these differences may affect the RLHF interpretation of well-known technical results in social choice.
format Preprint
id arxiv_https___arxiv_org_abs_2404_13038
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mapping Social Choice Theory to RLHF
Dai, Jessica
Fleisig, Eve
Artificial Intelligence
Computers and Society
Recent work on the limitations of using reinforcement learning from human feedback (RLHF) to incorporate human preferences into model behavior often raises social choice theory as a reference point. Social choice theory's analysis of settings such as voting mechanisms provides technical infrastructure that can inform how to aggregate human preferences amid disagreement. We analyze the problem settings of social choice and RLHF, identify key differences between them, and discuss how these differences may affect the RLHF interpretation of well-known technical results in social choice.
title Mapping Social Choice Theory to RLHF
topic Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2404.13038