PERSONA: A Reproducible Testbed for Pluralistic Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Castricato, Louis, Lile, Nathan, Rafailov, Rafael, Fränken, Jan-Philipp, Finn, Chelsea
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913443994402816
author Castricato, Louis
Lile, Nathan
Rafailov, Rafael
Fränken, Jan-Philipp
Finn, Chelsea
author_facet Castricato, Louis
Lile, Nathan
Rafailov, Rafael
Fränken, Jan-Philipp
Finn, Chelsea
contents The rapid advancement of language models (LMs) necessitates robust alignment with diverse user values. However, current preference optimization approaches often fail to capture the plurality of user opinions, instead reinforcing majority viewpoints and marginalizing minority perspectives. We introduce PERSONA, a reproducible test bed designed to evaluate and improve pluralistic alignment of LMs. We procedurally generate diverse user profiles from US census data, resulting in 1,586 synthetic personas with varied demographic and idiosyncratic attributes. We then generate a large-scale evaluation dataset containing 3,868 prompts and 317,200 feedback pairs obtained from our synthetic personas. Leveraging this dataset, we systematically evaluate LM capabilities in role-playing diverse users, verified through human judges, and the establishment of both a benchmark, PERSONA Bench, for pluralistic alignment approaches as well as an extensive dataset to create new and future benchmarks. The full dataset and benchmarks are available here: https://www.synthlabs.ai/research/persona.
format Preprint
id arxiv_https___arxiv_org_abs_2407_17387
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PERSONA: A Reproducible Testbed for Pluralistic Alignment
Castricato, Louis
Lile, Nathan
Rafailov, Rafael
Fränken, Jan-Philipp
Finn, Chelsea
Computation and Language
The rapid advancement of language models (LMs) necessitates robust alignment with diverse user values. However, current preference optimization approaches often fail to capture the plurality of user opinions, instead reinforcing majority viewpoints and marginalizing minority perspectives. We introduce PERSONA, a reproducible test bed designed to evaluate and improve pluralistic alignment of LMs. We procedurally generate diverse user profiles from US census data, resulting in 1,586 synthetic personas with varied demographic and idiosyncratic attributes. We then generate a large-scale evaluation dataset containing 3,868 prompts and 317,200 feedback pairs obtained from our synthetic personas. Leveraging this dataset, we systematically evaluate LM capabilities in role-playing diverse users, verified through human judges, and the establishment of both a benchmark, PERSONA Bench, for pluralistic alignment approaches as well as an extensive dataset to create new and future benchmarks. The full dataset and benchmarks are available here: https://www.synthlabs.ai/research/persona.
title PERSONA: A Reproducible Testbed for Pluralistic Alignment
topic Computation and Language
url https://arxiv.org/abs/2407.17387