ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ren, Yuanyi, Ye, Haoran, Fang, Hanjun, Zhang, Xin, Song, Guojie
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910475190534144
author Ren, Yuanyi
Ye, Haoran
Fang, Hanjun
Zhang, Xin
Song, Guojie
author_facet Ren, Yuanyi
Ye, Haoran
Fang, Hanjun
Zhang, Xin
Song, Guojie
contents Large Language Models (LLMs) are transforming diverse fields and gaining increasing influence as human proxies. This development underscores the urgent need for evaluating value orientations and understanding of LLMs to ensure their responsible integration into public-facing applications. This work introduces ValueBench, the first comprehensive psychometric benchmark for evaluating value orientations and value understanding in LLMs. ValueBench collects data from 44 established psychometric inventories, encompassing 453 multifaceted value dimensions. We propose an evaluation pipeline grounded in realistic human-AI interactions to probe value orientations, along with novel tasks for evaluating value understanding in an open-ended value space. With extensive experiments conducted on six representative LLMs, we unveil their shared and distinctive value orientations and exhibit their ability to approximate expert conclusions in value-related extraction and generation tasks. ValueBench is openly accessible at https://github.com/Value4AI/ValueBench.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04214
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models
Ren, Yuanyi
Ye, Haoran
Fang, Hanjun
Zhang, Xin
Song, Guojie
Computation and Language
Large Language Models (LLMs) are transforming diverse fields and gaining increasing influence as human proxies. This development underscores the urgent need for evaluating value orientations and understanding of LLMs to ensure their responsible integration into public-facing applications. This work introduces ValueBench, the first comprehensive psychometric benchmark for evaluating value orientations and value understanding in LLMs. ValueBench collects data from 44 established psychometric inventories, encompassing 453 multifaceted value dimensions. We propose an evaluation pipeline grounded in realistic human-AI interactions to probe value orientations, along with novel tasks for evaluating value understanding in an open-ended value space. With extensive experiments conducted on six representative LLMs, we unveil their shared and distinctive value orientations and exhibit their ability to approximate expert conclusions in value-related extraction and generation tasks. ValueBench is openly accessible at https://github.com/Value4AI/ValueBench.
title ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2406.04214