HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Zhaolu, Gong, Junhao, Yan, Jiaxu, Xia, Wanke, Wang, Yian, Wang, Ziwen, Ding, Huaxuan, Cheng, Zhuo, Cao, Wenhao, Feng, Zhiyuan, He, Siqi, Yan, Shannan, Chen, Junzhe, He, Xiaomin, Jiang, Chaoya, Ye, Wei, Yu, Kaidong, Li, Xuelong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912937841524736
author Kang, Zhaolu
Gong, Junhao
Yan, Jiaxu
Xia, Wanke
Wang, Yian
Wang, Ziwen
Ding, Huaxuan
Cheng, Zhuo
Cao, Wenhao
Feng, Zhiyuan
He, Siqi
Yan, Shannan
Chen, Junzhe
He, Xiaomin
Jiang, Chaoya
Ye, Wei
Yu, Kaidong
Li, Xuelong
author_facet Kang, Zhaolu
Gong, Junhao
Yan, Jiaxu
Xia, Wanke
Wang, Yian
Wang, Ziwen
Ding, Huaxuan
Cheng, Zhuo
Cao, Wenhao
Feng, Zhiyuan
He, Siqi
Yan, Shannan
Chen, Junzhe
He, Xiaomin
Jiang, Chaoya
Ye, Wei
Yu, Kaidong
Li, Xuelong
contents Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchmarks for evaluating MLLMs primarily emphasize general knowledge and vertical step-by-step reasoning typical of STEM disciplines, while overlooking the distinct needs and potential of the Humanities and Social Sciences (HSS). Tasks in the HSS domain require more horizontal, interdisciplinary thinking and a deep integration of knowledge across related fields, which presents unique challenges for MLLMs, particularly in linking abstract concepts with corresponding visual representations. Addressing this gap, we present HSSBench, a dedicated benchmark designed to assess the capabilities of MLLMs on HSS tasks in multiple languages, including the six official languages of the United Nations. We also introduce a novel data generation pipeline tailored for HSS scenarios, in which multiple domain experts and automated agents collaborate to generate and iteratively refine each sample. HSSBench contains over 13,000 meticulously designed samples, covering six key categories. We benchmark more than 20 mainstream MLLMs on HSSBench and demonstrate that it poses significant challenges even for state-of-the-art models. We hope that this benchmark will inspire further research into enhancing the cross-disciplinary reasoning abilities of MLLMs, especially their capacity to internalize and connect knowledge across fields.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03922
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models
Kang, Zhaolu
Gong, Junhao
Yan, Jiaxu
Xia, Wanke
Wang, Yian
Wang, Ziwen
Ding, Huaxuan
Cheng, Zhuo
Cao, Wenhao
Feng, Zhiyuan
He, Siqi
Yan, Shannan
Chen, Junzhe
He, Xiaomin
Jiang, Chaoya
Ye, Wei
Yu, Kaidong
Li, Xuelong
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchmarks for evaluating MLLMs primarily emphasize general knowledge and vertical step-by-step reasoning typical of STEM disciplines, while overlooking the distinct needs and potential of the Humanities and Social Sciences (HSS). Tasks in the HSS domain require more horizontal, interdisciplinary thinking and a deep integration of knowledge across related fields, which presents unique challenges for MLLMs, particularly in linking abstract concepts with corresponding visual representations. Addressing this gap, we present HSSBench, a dedicated benchmark designed to assess the capabilities of MLLMs on HSS tasks in multiple languages, including the six official languages of the United Nations. We also introduce a novel data generation pipeline tailored for HSS scenarios, in which multiple domain experts and automated agents collaborate to generate and iteratively refine each sample. HSSBench contains over 13,000 meticulously designed samples, covering six key categories. We benchmark more than 20 mainstream MLLMs on HSSBench and demonstrate that it poses significant challenges even for state-of-the-art models. We hope that this benchmark will inspire further research into enhancing the cross-disciplinary reasoning abilities of MLLMs, especially their capacity to internalize and connect knowledge across fields.
title HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.03922