Heterogeneous Value Alignment Evaluation for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhaowei, Zhang, Ceyao, Liu, Nian, Qi, Siyuan, Rong, Ziqi, Zhu, Song-Chun, Cui, Shuguang, Yang, Yaodong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929206934372352
author Zhang, Zhaowei
Zhang, Ceyao
Liu, Nian
Qi, Siyuan
Rong, Ziqi
Zhu, Song-Chun
Cui, Shuguang
Yang, Yaodong
author_facet Zhang, Zhaowei
Zhang, Ceyao
Liu, Nian
Qi, Siyuan
Rong, Ziqi
Zhu, Song-Chun
Cui, Shuguang
Yang, Yaodong
contents The emergent capabilities of Large Language Models (LLMs) have made it crucial to align their values with those of humans. However, current methodologies typically attempt to assign value as an attribute to LLMs, yet lack attention to the ability to pursue value and the importance of transferring heterogeneous values in specific practical applications. In this paper, we propose a Heterogeneous Value Alignment Evaluation (HVAE) system, designed to assess the success of aligning LLMs with heterogeneous values. Specifically, our approach first brings the Social Value Orientation (SVO) framework from social psychology, which corresponds to how much weight a person attaches to the welfare of others in relation to their own. We then assign the LLMs with different social values and measure whether their behaviors align with the inducing values. We conduct evaluations with new auto-metric \textit{value rationality} to represent the ability of LLMs to align with specific values. Evaluating the value rationality of five mainstream LLMs, we discern a propensity in LLMs towards neutral values over pronounced personal values. By examining the behavior of these LLMs, we contribute to a deeper insight into the value alignment of LLMs within a heterogeneous value system.
format Preprint
id arxiv_https___arxiv_org_abs_2305_17147
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Heterogeneous Value Alignment Evaluation for Large Language Models
Zhang, Zhaowei
Zhang, Ceyao
Liu, Nian
Qi, Siyuan
Rong, Ziqi
Zhu, Song-Chun
Cui, Shuguang
Yang, Yaodong
Computation and Language
Artificial Intelligence
Human-Computer Interaction
Machine Learning
The emergent capabilities of Large Language Models (LLMs) have made it crucial to align their values with those of humans. However, current methodologies typically attempt to assign value as an attribute to LLMs, yet lack attention to the ability to pursue value and the importance of transferring heterogeneous values in specific practical applications. In this paper, we propose a Heterogeneous Value Alignment Evaluation (HVAE) system, designed to assess the success of aligning LLMs with heterogeneous values. Specifically, our approach first brings the Social Value Orientation (SVO) framework from social psychology, which corresponds to how much weight a person attaches to the welfare of others in relation to their own. We then assign the LLMs with different social values and measure whether their behaviors align with the inducing values. We conduct evaluations with new auto-metric \textit{value rationality} to represent the ability of LLMs to align with specific values. Evaluating the value rationality of five mainstream LLMs, we discern a propensity in LLMs towards neutral values over pronounced personal values. By examining the behavior of these LLMs, we contribute to a deeper insight into the value alignment of LLMs within a heterogeneous value system.
title Heterogeneous Value Alignment Evaluation for Large Language Models
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2305.17147