HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Keliang, Yang, Zaifei, Zhao, Jiahe, Shen, Hongze, Hou, Ruibing, Chang, Hong, Shan, Shiguang, Chen, Xilin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913539136946176
author Li, Keliang
Yang, Zaifei
Zhao, Jiahe
Shen, Hongze
Hou, Ruibing
Chang, Hong
Shan, Shiguang
Chen, Xilin
author_facet Li, Keliang
Yang, Zaifei
Zhao, Jiahe
Shen, Hongze
Hou, Ruibing
Chang, Hong
Shan, Shiguang
Chen, Xilin
contents The significant advancements in visual understanding and instruction following from Multimodal Large Language Models (MLLMs) have opened up more possibilities for broader applications in diverse and universal human-centric scenarios. However, existing image-text data may not support the precise modality alignment and integration of multi-grained information, which is crucial for human-centric visual understanding. In this paper, we introduce HERM-Bench, a benchmark for evaluating the human-centric understanding capabilities of MLLMs. Our work reveals the limitations of existing MLLMs in understanding complex human-centric scenarios. To address these challenges, we present HERM-100K, a comprehensive dataset with multi-level human-centric annotations, aimed at enhancing MLLMs' training. Furthermore, we develop HERM-7B, a MLLM that leverages enhanced training data from HERM-100K. Evaluations on HERM-Bench demonstrate that HERM-7B significantly outperforms existing MLLMs across various human-centric dimensions, reflecting the current inadequacy of data annotations used in MLLM training for human-centric visual understanding. This research emphasizes the importance of specialized datasets and benchmarks in advancing the MLLMs' capabilities for human-centric understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06777
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
Li, Keliang
Yang, Zaifei
Zhao, Jiahe
Shen, Hongze
Hou, Ruibing
Chang, Hong
Shan, Shiguang
Chen, Xilin
Computer Vision and Pattern Recognition
The significant advancements in visual understanding and instruction following from Multimodal Large Language Models (MLLMs) have opened up more possibilities for broader applications in diverse and universal human-centric scenarios. However, existing image-text data may not support the precise modality alignment and integration of multi-grained information, which is crucial for human-centric visual understanding. In this paper, we introduce HERM-Bench, a benchmark for evaluating the human-centric understanding capabilities of MLLMs. Our work reveals the limitations of existing MLLMs in understanding complex human-centric scenarios. To address these challenges, we present HERM-100K, a comprehensive dataset with multi-level human-centric annotations, aimed at enhancing MLLMs' training. Furthermore, we develop HERM-7B, a MLLM that leverages enhanced training data from HERM-100K. Evaluations on HERM-Bench demonstrate that HERM-7B significantly outperforms existing MLLMs across various human-centric dimensions, reflecting the current inadequacy of data annotations used in MLLM training for human-centric visual understanding. This research emphasizes the importance of specialized datasets and benchmarks in advancing the MLLMs' capabilities for human-centric understanding.
title HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.06777