A Roadmap to Pluralistic Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sorensen, Taylor, Moore, Jared, Fisher, Jillian, Gordon, Mitchell, Mireshghallah, Niloofar, Rytting, Christopher Michael, Ye, Andre, Jiang, Liwei, Lu, Ximing, Dziri, Nouha, Althoff, Tim, Choi, Yejin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909291893489664
author Sorensen, Taylor
Moore, Jared
Fisher, Jillian
Gordon, Mitchell
Mireshghallah, Niloofar
Rytting, Christopher Michael
Ye, Andre
Jiang, Liwei
Lu, Ximing
Dziri, Nouha
Althoff, Tim
Choi, Yejin
author_facet Sorensen, Taylor
Moore, Jared
Fisher, Jillian
Gordon, Mitchell
Mireshghallah, Niloofar
Rytting, Christopher Michael
Ye, Andre
Jiang, Liwei
Lu, Ximing
Dziri, Nouha
Althoff, Tim
Choi, Yejin
contents With increased power and prevalence of AI systems, it is ever more critical that AI systems are designed to serve all, i.e., people with diverse values and perspectives. However, aligning models to serve pluralistic human values remains an open research question. In this piece, we propose a roadmap to pluralistic alignment, specifically using language models as a test bed. We identify and formalize three possible ways to define and operationalize pluralism in AI systems: 1) Overton pluralistic models that present a spectrum of reasonable responses; 2) Steerably pluralistic models that can steer to reflect certain perspectives; and 3) Distributionally pluralistic models that are well-calibrated to a given population in distribution. We also formalize and discuss three possible classes of pluralistic benchmarks: 1) Multi-objective benchmarks, 2) Trade-off steerable benchmarks, which incentivize models to steer to arbitrary trade-offs, and 3) Jury-pluralistic benchmarks which explicitly model diverse human ratings. We use this framework to argue that current alignment techniques may be fundamentally limited for pluralistic AI; indeed, we highlight empirical evidence, both from our own experiments and from other work, that standard alignment procedures might reduce distributional pluralism in models, motivating the need for further research on pluralistic alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2402_05070
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Roadmap to Pluralistic Alignment
Sorensen, Taylor
Moore, Jared
Fisher, Jillian
Gordon, Mitchell
Mireshghallah, Niloofar
Rytting, Christopher Michael
Ye, Andre
Jiang, Liwei
Lu, Ximing
Dziri, Nouha
Althoff, Tim
Choi, Yejin
Artificial Intelligence
Computation and Language
Information Retrieval
With increased power and prevalence of AI systems, it is ever more critical that AI systems are designed to serve all, i.e., people with diverse values and perspectives. However, aligning models to serve pluralistic human values remains an open research question. In this piece, we propose a roadmap to pluralistic alignment, specifically using language models as a test bed. We identify and formalize three possible ways to define and operationalize pluralism in AI systems: 1) Overton pluralistic models that present a spectrum of reasonable responses; 2) Steerably pluralistic models that can steer to reflect certain perspectives; and 3) Distributionally pluralistic models that are well-calibrated to a given population in distribution. We also formalize and discuss three possible classes of pluralistic benchmarks: 1) Multi-objective benchmarks, 2) Trade-off steerable benchmarks, which incentivize models to steer to arbitrary trade-offs, and 3) Jury-pluralistic benchmarks which explicitly model diverse human ratings. We use this framework to argue that current alignment techniques may be fundamentally limited for pluralistic AI; indeed, we highlight empirical evidence, both from our own experiments and from other work, that standard alignment procedures might reduce distributional pluralism in models, motivating the need for further research on pluralistic alignment.
title A Roadmap to Pluralistic Alignment
topic Artificial Intelligence
Computation and Language
Information Retrieval
url https://arxiv.org/abs/2402.05070