Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yuansen, Tang, Haiming, Peng, Jinlong, Zhang, Jiangning, Ji, Xiaozhong, He, Qingdong, Wu, Wenbin, Luo, Donghao, Gan, Zhenye, Zhu, Junwei, Shen, Yunhang, Fu, Chaoyou, Wang, Chengjie, Hu, Xiaobin, Yan, Shuicheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914094860206080
author Liu, Yuansen
Tang, Haiming
Peng, Jinlong
Zhang, Jiangning
Ji, Xiaozhong
He, Qingdong
Wu, Wenbin
Luo, Donghao
Gan, Zhenye
Zhu, Junwei
Shen, Yunhang
Fu, Chaoyou
Wang, Chengjie
Hu, Xiaobin
Yan, Shuicheng
author_facet Liu, Yuansen
Tang, Haiming
Peng, Jinlong
Zhang, Jiangning
Ji, Xiaozhong
He, Qingdong
Wu, Wenbin
Luo, Donghao
Gan, Zhenye
Zhu, Junwei
Shen, Yunhang
Fu, Chaoyou
Wang, Chengjie
Hu, Xiaobin
Yan, Shuicheng
contents Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks. However, their capacity to comprehend human-centric scenes has rarely been explored, primarily due to the absence of comprehensive evaluation benchmarks that take into account both the human-oriented granular level and higher-dimensional causal reasoning ability. Such high-quality evaluation benchmarks face tough obstacles, given the physical complexity of the human body and the difficulty of annotating granular structures. In this paper, we propose Human-MME, a curated benchmark designed to provide a more holistic evaluation of MLLMs in human-centric scene understanding. Compared with other existing benchmarks, our work provides three key features: 1. Diversity in human scene, spanning 4 primary visual domains with 15 secondary domains and 43 sub-fields to ensure broad scenario coverage. 2. Progressive and diverse evaluation dimensions, evaluating the human-based activities progressively from the human-oriented granular perception to the higher-dimensional reasoning, consisting of eight dimensions with 19,945 real-world image question pairs and an evaluation suite. 3. High-quality annotations with rich data paradigms, constructing the automated annotation pipeline and human-annotation platform, supporting rigorous manual labeling to facilitate precise and reliable model assessment. Our benchmark extends the single-target understanding to the multi-person and multi-image mutual understanding by constructing the choice, short-answer, grounding, ranking and judgment question components, and complex questions of their combination. The extensive experiments on 17 state-of-the-art MLLMs effectively expose the limitations and guide future MLLMs research toward better human-centric image understanding. All data and code are available at https://github.com/Yuan-Hou/Human-MME.
format Preprint
id arxiv_https___arxiv_org_abs_2509_26165
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
Liu, Yuansen
Tang, Haiming
Peng, Jinlong
Zhang, Jiangning
Ji, Xiaozhong
He, Qingdong
Wu, Wenbin
Luo, Donghao
Gan, Zhenye
Zhu, Junwei
Shen, Yunhang
Fu, Chaoyou
Wang, Chengjie
Hu, Xiaobin
Yan, Shuicheng
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks. However, their capacity to comprehend human-centric scenes has rarely been explored, primarily due to the absence of comprehensive evaluation benchmarks that take into account both the human-oriented granular level and higher-dimensional causal reasoning ability. Such high-quality evaluation benchmarks face tough obstacles, given the physical complexity of the human body and the difficulty of annotating granular structures. In this paper, we propose Human-MME, a curated benchmark designed to provide a more holistic evaluation of MLLMs in human-centric scene understanding. Compared with other existing benchmarks, our work provides three key features: 1. Diversity in human scene, spanning 4 primary visual domains with 15 secondary domains and 43 sub-fields to ensure broad scenario coverage. 2. Progressive and diverse evaluation dimensions, evaluating the human-based activities progressively from the human-oriented granular perception to the higher-dimensional reasoning, consisting of eight dimensions with 19,945 real-world image question pairs and an evaluation suite. 3. High-quality annotations with rich data paradigms, constructing the automated annotation pipeline and human-annotation platform, supporting rigorous manual labeling to facilitate precise and reliable model assessment. Our benchmark extends the single-target understanding to the multi-person and multi-image mutual understanding by constructing the choice, short-answer, grounding, ranking and judgment question components, and complex questions of their combination. The extensive experiments on 17 state-of-the-art MLLMs effectively expose the limitations and guide future MLLMs research toward better human-centric image understanding. All data and code are available at https://github.com/Yuan-Hou/Human-MME.
title Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.26165