HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cai, Yuxuan, Zhang, Jiangning, Gan, Zhenye, He, Qingdong, Hu, Xiaobin, Zhu, Junwei, Wang, Yabiao, Wang, Chengjie, Xue, Zhucun, Fu, Chaoyou, He, Xinwei, Bai, Xiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914065995005952
author Cai, Yuxuan
Zhang, Jiangning
Gan, Zhenye
He, Qingdong
Hu, Xiaobin
Zhu, Junwei
Wang, Yabiao
Wang, Chengjie
Xue, Zhucun
Fu, Chaoyou
He, Xinwei
Bai, Xiang
author_facet Cai, Yuxuan
Zhang, Jiangning
Gan, Zhenye
He, Qingdong
Hu, Xiaobin
Zhu, Junwei
Wang, Yabiao
Wang, Chengjie
Xue, Zhucun
Fu, Chaoyou
He, Xinwei
Bai, Xiang
contents Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks involving both images and videos. However, their capacity to comprehend human-centric video data remains underexplored, primarily due to the absence of comprehensive and high-quality evaluation benchmarks. Existing human-centric benchmarks predominantly emphasize video generation quality and action recognition, while overlooking essential perceptual and cognitive abilities required in human-centered scenarios. Furthermore, they are often limited by single-question paradigms and overly simplistic evaluation metrics. To address above limitations, we propose a modern HV-MMBench, a rigorously curated benchmark designed to provide a more holistic evaluation of MLLMs in human-centric video understanding. Compared to existing human-centric video benchmarks, our work offers the following key features: (1) Diverse evaluation dimensions: HV-MMBench encompasses 13 tasks, ranging from basic attribute perception (e.g., age estimation, emotion recognition) to advanced cognitive reasoning (e.g., social relationship prediction, intention prediction), enabling comprehensive assessment of model capabilities; (2) Varied data types: The benchmark includes multiple-choice, fill-in-blank, true/false, and open-ended question formats, combined with diverse evaluation metrics, to more accurately and robustly reflect model performance; (3) Multi-domain video coverage: The benchmark spans 50 distinct visual scenarios, enabling comprehensive evaluation across fine-grained scene variations; (4) Temporal coverage: The benchmark covers videos from short-term (10 seconds) to long-term (up to 30min) durations, supporting systematic analysis of models temporal reasoning abilities across diverse contextual lengths.
format Preprint
id arxiv_https___arxiv_org_abs_2507_04909
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
Cai, Yuxuan
Zhang, Jiangning
Gan, Zhenye
He, Qingdong
Hu, Xiaobin
Zhu, Junwei
Wang, Yabiao
Wang, Chengjie
Xue, Zhucun
Fu, Chaoyou
He, Xinwei
Bai, Xiang
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks involving both images and videos. However, their capacity to comprehend human-centric video data remains underexplored, primarily due to the absence of comprehensive and high-quality evaluation benchmarks. Existing human-centric benchmarks predominantly emphasize video generation quality and action recognition, while overlooking essential perceptual and cognitive abilities required in human-centered scenarios. Furthermore, they are often limited by single-question paradigms and overly simplistic evaluation metrics. To address above limitations, we propose a modern HV-MMBench, a rigorously curated benchmark designed to provide a more holistic evaluation of MLLMs in human-centric video understanding. Compared to existing human-centric video benchmarks, our work offers the following key features: (1) Diverse evaluation dimensions: HV-MMBench encompasses 13 tasks, ranging from basic attribute perception (e.g., age estimation, emotion recognition) to advanced cognitive reasoning (e.g., social relationship prediction, intention prediction), enabling comprehensive assessment of model capabilities; (2) Varied data types: The benchmark includes multiple-choice, fill-in-blank, true/false, and open-ended question formats, combined with diverse evaluation metrics, to more accurately and robustly reflect model performance; (3) Multi-domain video coverage: The benchmark spans 50 distinct visual scenarios, enabling comprehensive evaluation across fine-grained scene variations; (4) Temporal coverage: The benchmark covers videos from short-term (10 seconds) to long-term (up to 30min) durations, supporting systematic analysis of models temporal reasoning abilities across diverse contextual lengths.
title HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2507.04909