Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Zhengliang, Ma, Ruotian, Huang, Jen-tse, Ma, Xinbei, Chen, Xingyu, Wang, Mengru, Yang, Qu, Wang, Yue, Ye, Fanghua, Chen, Ziyang, Wang, Shanyi, Li, Cixing, Wang, Wenxuan, Tu, Zhaopeng, Li, Xiaolong, Ren, Zhaochun, Linus
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912621162135552
author Shi, Zhengliang
Ma, Ruotian
Huang, Jen-tse
Ma, Xinbei
Chen, Xingyu
Wang, Mengru
Yang, Qu
Wang, Yue
Ye, Fanghua
Chen, Ziyang
Wang, Shanyi
Li, Cixing
Wang, Wenxuan
Tu, Zhaopeng
Li, Xiaolong
Ren, Zhaochun
Linus
author_facet Shi, Zhengliang
Ma, Ruotian
Huang, Jen-tse
Ma, Xinbei
Chen, Xingyu
Wang, Mengru
Yang, Qu
Wang, Yue
Ye, Fanghua
Chen, Ziyang
Wang, Shanyi
Li, Cixing
Wang, Wenxuan
Tu, Zhaopeng
Li, Xiaolong
Ren, Zhaochun
Linus
contents Large language models (LLMs) are increasingly entrusted with high-stakes decisions that affect human welfare. However, the principles and values that guide these models when distributing scarce societal resources remain largely unexamined. To address this, we introduce the Social Welfare Function (SWF) Benchmark, a dynamic simulation environment where an LLM acts as a sovereign allocator, distributing tasks to a heterogeneous community of recipients. The benchmark is designed to create a persistent trade-off between maximizing collective efficiency (measured by Return on Investment) and ensuring distributive fairness (measured by the Gini coefficient). We evaluate 20 state-of-the-art LLMs and present the first leaderboard for social welfare allocation. Our findings reveal three key insights: (i) A model's general conversational ability, as measured by popular leaderboards, is a poor predictor of its allocation skill. (ii) Most LLMs exhibit a strong default utilitarian orientation, prioritizing group productivity at the expense of severe inequality. (iii) Allocation strategies are highly vulnerable, easily perturbed by output-length constraints and social-influence framing. These results highlight the risks of deploying current LLMs as societal decision-makers and underscore the need for specialized benchmarks and targeted alignment for AI governance.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01164
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare
Shi, Zhengliang
Ma, Ruotian
Huang, Jen-tse
Ma, Xinbei
Chen, Xingyu
Wang, Mengru
Yang, Qu
Wang, Yue
Ye, Fanghua
Chen, Ziyang
Wang, Shanyi
Li, Cixing
Wang, Wenxuan
Tu, Zhaopeng
Li, Xiaolong
Ren, Zhaochun
Linus
Computation and Language
Artificial Intelligence
Computers and Society
Human-Computer Interaction
Large language models (LLMs) are increasingly entrusted with high-stakes decisions that affect human welfare. However, the principles and values that guide these models when distributing scarce societal resources remain largely unexamined. To address this, we introduce the Social Welfare Function (SWF) Benchmark, a dynamic simulation environment where an LLM acts as a sovereign allocator, distributing tasks to a heterogeneous community of recipients. The benchmark is designed to create a persistent trade-off between maximizing collective efficiency (measured by Return on Investment) and ensuring distributive fairness (measured by the Gini coefficient). We evaluate 20 state-of-the-art LLMs and present the first leaderboard for social welfare allocation. Our findings reveal three key insights: (i) A model's general conversational ability, as measured by popular leaderboards, is a poor predictor of its allocation skill. (ii) Most LLMs exhibit a strong default utilitarian orientation, prioritizing group productivity at the expense of severe inequality. (iii) Allocation strategies are highly vulnerable, easily perturbed by output-length constraints and social-influence framing. These results highlight the risks of deploying current LLMs as societal decision-makers and underscore the need for specialized benchmarks and targeted alignment for AI governance.
title Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare
topic Computation and Language
Artificial Intelligence
Computers and Society
Human-Computer Interaction
url https://arxiv.org/abs/2510.01164