HeartBench: Probing Core Dimensions of Anthropomorphic Intelligence in LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Jiaxin, Tu, Peiyi, Chen, Wenyu, Zhuang, Yihong, Ling, Xinxia, Zhou, Anji, Wang, Chenxi, Han, Zhuo, Yang, Zhengkai, Zhao, Junbo, Huang, Zenan, Wang, Yuanyuan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909976078843904
author Liu, Jiaxin
Tu, Peiyi
Chen, Wenyu
Zhuang, Yihong
Ling, Xinxia
Zhou, Anji
Wang, Chenxi
Han, Zhuo
Yang, Zhengkai
Zhao, Junbo
Huang, Zenan
Wang, Yuanyuan
author_facet Liu, Jiaxin
Tu, Peiyi
Chen, Wenyu
Zhuang, Yihong
Ling, Xinxia
Zhou, Anji
Wang, Chenxi
Han, Zhuo
Yang, Zhengkai
Zhao, Junbo
Huang, Zenan
Wang, Yuanyuan
contents While Large Language Models (LLMs) have achieved remarkable success in cognitive and reasoning benchmarks, they exhibit a persistent deficit in anthropomorphic intelligence-the capacity to navigate complex social, emotional, and ethical nuances. This gap is particularly acute in the Chinese linguistic and cultural context, where a lack of specialized evaluation frameworks and high-quality socio-emotional data impedes progress. To address these limitations, we present HeartBench, a framework designed to evaluate the integrated emotional, cultural, and ethical dimensions of Chinese LLMs. Grounded in authentic psychological counseling scenarios and developed in collaboration with clinical experts, the benchmark is structured around a theory-driven taxonomy comprising five primary dimensions and 15 secondary capabilities. We implement a case-specific, rubric-based methodology that translates abstract human-like traits into granular, measurable criteria through a ``reasoning-before-scoring'' evaluation protocol. Our assessment of 13 state-of-the-art LLMs indicates a substantial performance ceiling: even leading models achieve only 60% of the expert-defined ideal score. Furthermore, analysis using a difficulty-stratified ``Hard Set'' reveals a significant performance decay in scenarios involving subtle emotional subtexts and complex ethical trade-offs. HeartBench establishes a standardized metric for anthropomorphic AI evaluation and provides a methodological blueprint for constructing high-quality, human-aligned training data.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21849
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HeartBench: Probing Core Dimensions of Anthropomorphic Intelligence in LLMs
Liu, Jiaxin
Tu, Peiyi
Chen, Wenyu
Zhuang, Yihong
Ling, Xinxia
Zhou, Anji
Wang, Chenxi
Han, Zhuo
Yang, Zhengkai
Zhao, Junbo
Huang, Zenan
Wang, Yuanyuan
Computation and Language
Artificial Intelligence
While Large Language Models (LLMs) have achieved remarkable success in cognitive and reasoning benchmarks, they exhibit a persistent deficit in anthropomorphic intelligence-the capacity to navigate complex social, emotional, and ethical nuances. This gap is particularly acute in the Chinese linguistic and cultural context, where a lack of specialized evaluation frameworks and high-quality socio-emotional data impedes progress. To address these limitations, we present HeartBench, a framework designed to evaluate the integrated emotional, cultural, and ethical dimensions of Chinese LLMs. Grounded in authentic psychological counseling scenarios and developed in collaboration with clinical experts, the benchmark is structured around a theory-driven taxonomy comprising five primary dimensions and 15 secondary capabilities. We implement a case-specific, rubric-based methodology that translates abstract human-like traits into granular, measurable criteria through a ``reasoning-before-scoring'' evaluation protocol. Our assessment of 13 state-of-the-art LLMs indicates a substantial performance ceiling: even leading models achieve only 60% of the expert-defined ideal score. Furthermore, analysis using a difficulty-stratified ``Hard Set'' reveals a significant performance decay in scenarios involving subtle emotional subtexts and complex ethical trade-offs. HeartBench establishes a standardized metric for anthropomorphic AI evaluation and provides a methodological blueprint for constructing high-quality, human-aligned training data.
title HeartBench: Probing Core Dimensions of Anthropomorphic Intelligence in LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.21849