MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yong, Xixian, Lian, Jianxun, Yi, Xiaoyuan, Zhou, Xiao, Xie, Xing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909649987436544
author Yong, Xixian
Lian, Jianxun
Yi, Xiaoyuan
Zhou, Xiao
Xie, Xing
author_facet Yong, Xixian
Lian, Jianxun
Yi, Xiaoyuan
Zhou, Xiao
Xie, Xing
contents Large language models (LLMs) have been widely adopted as the core of agent frameworks in various scenarios, such as social simulations and AI companions. However, the extent to which they can replicate human-like motivations remains an underexplored question. Existing benchmarks are constrained by simplistic scenarios and the absence of character identities, resulting in an information asymmetry with real-world situations. To address this gap, we propose MotiveBench, which consists of 200 rich contextual scenarios and 600 reasoning tasks covering multiple levels of motivation. Using MotiveBench, we conduct extensive experiments on seven popular model families, comparing different scales and versions within each family. The results show that even the most advanced LLMs still fall short in achieving human-like motivational reasoning. Our analysis reveals key findings, including the difficulty LLMs face in reasoning about "love & belonging" motivations and their tendency toward excessive rationality and idealism. These insights highlight a promising direction for future research on the humanization of LLMs. The dataset, benchmark, and code are available at https://aka.ms/motivebench.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13065
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
Yong, Xixian
Lian, Jianxun
Yi, Xiaoyuan
Zhou, Xiao
Xie, Xing
Computation and Language
Artificial Intelligence
Large language models (LLMs) have been widely adopted as the core of agent frameworks in various scenarios, such as social simulations and AI companions. However, the extent to which they can replicate human-like motivations remains an underexplored question. Existing benchmarks are constrained by simplistic scenarios and the absence of character identities, resulting in an information asymmetry with real-world situations. To address this gap, we propose MotiveBench, which consists of 200 rich contextual scenarios and 600 reasoning tasks covering multiple levels of motivation. Using MotiveBench, we conduct extensive experiments on seven popular model families, comparing different scales and versions within each family. The results show that even the most advanced LLMs still fall short in achieving human-like motivational reasoning. Our analysis reveals key findings, including the difficulty LLMs face in reasoning about "love & belonging" motivations and their tendency toward excessive rationality and idealism. These insights highlight a promising direction for future research on the humanization of LLMs. The dataset, benchmark, and code are available at https://aka.ms/motivebench.
title MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.13065