Can Language Models Reason about Individualistic Human Values and Preferences?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Liwei, Sorensen, Taylor, Levine, Sydney, Choi, Yejin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913868327944192
author Jiang, Liwei
Sorensen, Taylor
Levine, Sydney
Choi, Yejin
author_facet Jiang, Liwei
Sorensen, Taylor
Levine, Sydney
Choi, Yejin
contents Recent calls for pluralistic alignment emphasize that AI systems should address the diverse needs of all people. Yet, efforts in this space often require sorting people into fixed buckets of pre-specified diversity-defining dimensions (e.g., demographics), risking smoothing out individualistic variations or even stereotyping. To achieve an authentic representation of diversity that respects individuality, we propose individualistic alignment. While individualistic alignment can take various forms, we introduce IndieValueCatalog, a dataset transformed from the influential World Values Survey (WVS), to study language models (LMs) on the specific challenge of individualistic value reasoning. Given a sample of an individual's value-expressing statements, models are tasked with predicting this person's value judgments in novel cases. With IndieValueCatalog, we reveal critical limitations in frontier LMs, which achieve only 55 % to 65% accuracy in predicting individualistic values. Moreover, our results highlight that a precise description of individualistic values cannot be approximated only with demographic information. We also identify a partiality of LMs in reasoning about global individualistic values, as measured by our proposed Value Inequity Index (σInequity). Finally, we train a series of IndieValueReasoners to reveal new patterns and dynamics into global human values.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03868
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Can Language Models Reason about Individualistic Human Values and Preferences?
Jiang, Liwei
Sorensen, Taylor
Levine, Sydney
Choi, Yejin
Computation and Language
Recent calls for pluralistic alignment emphasize that AI systems should address the diverse needs of all people. Yet, efforts in this space often require sorting people into fixed buckets of pre-specified diversity-defining dimensions (e.g., demographics), risking smoothing out individualistic variations or even stereotyping. To achieve an authentic representation of diversity that respects individuality, we propose individualistic alignment. While individualistic alignment can take various forms, we introduce IndieValueCatalog, a dataset transformed from the influential World Values Survey (WVS), to study language models (LMs) on the specific challenge of individualistic value reasoning. Given a sample of an individual's value-expressing statements, models are tasked with predicting this person's value judgments in novel cases. With IndieValueCatalog, we reveal critical limitations in frontier LMs, which achieve only 55 % to 65% accuracy in predicting individualistic values. Moreover, our results highlight that a precise description of individualistic values cannot be approximated only with demographic information. We also identify a partiality of LMs in reasoning about global individualistic values, as measured by our proposed Value Inequity Index (σInequity). Finally, we train a series of IndieValueReasoners to reveal new patterns and dynamics into global human values.
title Can Language Models Reason about Individualistic Human Values and Preferences?
topic Computation and Language
url https://arxiv.org/abs/2410.03868