Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912621162135552 |
|---|---|
| author | Shi, Zhengliang Ma, Ruotian Huang, Jen-tse Ma, Xinbei Chen, Xingyu Wang, Mengru Yang, Qu Wang, Yue Ye, Fanghua Chen, Ziyang Wang, Shanyi Li, Cixing Wang, Wenxuan Tu, Zhaopeng Li, Xiaolong Ren, Zhaochun Linus |
| author_facet | Shi, Zhengliang Ma, Ruotian Huang, Jen-tse Ma, Xinbei Chen, Xingyu Wang, Mengru Yang, Qu Wang, Yue Ye, Fanghua Chen, Ziyang Wang, Shanyi Li, Cixing Wang, Wenxuan Tu, Zhaopeng Li, Xiaolong Ren, Zhaochun Linus |
| contents | Large language models (LLMs) are increasingly entrusted with high-stakes decisions that affect human welfare. However, the principles and values that guide these models when distributing scarce societal resources remain largely unexamined. To address this, we introduce the Social Welfare Function (SWF) Benchmark, a dynamic simulation environment where an LLM acts as a sovereign allocator, distributing tasks to a heterogeneous community of recipients. The benchmark is designed to create a persistent trade-off between maximizing collective efficiency (measured by Return on Investment) and ensuring distributive fairness (measured by the Gini coefficient). We evaluate 20 state-of-the-art LLMs and present the first leaderboard for social welfare allocation. Our findings reveal three key insights: (i) A model's general conversational ability, as measured by popular leaderboards, is a poor predictor of its allocation skill. (ii) Most LLMs exhibit a strong default utilitarian orientation, prioritizing group productivity at the expense of severe inequality. (iii) Allocation strategies are highly vulnerable, easily perturbed by output-length constraints and social-influence framing. These results highlight the risks of deploying current LLMs as societal decision-makers and underscore the need for specialized benchmarks and targeted alignment for AI governance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_01164 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare Shi, Zhengliang Ma, Ruotian Huang, Jen-tse Ma, Xinbei Chen, Xingyu Wang, Mengru Yang, Qu Wang, Yue Ye, Fanghua Chen, Ziyang Wang, Shanyi Li, Cixing Wang, Wenxuan Tu, Zhaopeng Li, Xiaolong Ren, Zhaochun Linus Computation and Language Artificial Intelligence Computers and Society Human-Computer Interaction Large language models (LLMs) are increasingly entrusted with high-stakes decisions that affect human welfare. However, the principles and values that guide these models when distributing scarce societal resources remain largely unexamined. To address this, we introduce the Social Welfare Function (SWF) Benchmark, a dynamic simulation environment where an LLM acts as a sovereign allocator, distributing tasks to a heterogeneous community of recipients. The benchmark is designed to create a persistent trade-off between maximizing collective efficiency (measured by Return on Investment) and ensuring distributive fairness (measured by the Gini coefficient). We evaluate 20 state-of-the-art LLMs and present the first leaderboard for social welfare allocation. Our findings reveal three key insights: (i) A model's general conversational ability, as measured by popular leaderboards, is a poor predictor of its allocation skill. (ii) Most LLMs exhibit a strong default utilitarian orientation, prioritizing group productivity at the expense of severe inequality. (iii) Allocation strategies are highly vulnerable, easily perturbed by output-length constraints and social-influence framing. These results highlight the risks of deploying current LLMs as societal decision-makers and underscore the need for specialized benchmarks and targeted alignment for AI governance. |
| title | Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare |
| topic | Computation and Language Artificial Intelligence Computers and Society Human-Computer Interaction |
| url | https://arxiv.org/abs/2510.01164 |