MA-Bench: Towards Fine-grained Micro-Action Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Kun, Gu, Jihao, Wang, Fei, Wu, Zhiliang, Fan, Hehe, Guo, Dan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911548403875840
author Li, Kun
Gu, Jihao
Wang, Fei
Wu, Zhiliang
Fan, Hehe
Guo, Dan
author_facet Li, Kun
Gu, Jihao
Wang, Fei
Wu, Zhiliang
Fan, Hehe
Guo, Dan
contents With the rapid development of Multimodal Large Language Models (MLLMs), their potential in Micro-Action understanding, a vital role in human emotion analysis, remains unexplored due to the absence of specialized benchmarks. To tackle this issue, we present MA-Bench, a benchmark comprising 1,000 videos and a three-tier evaluation architecture that progressively examines micro-action perception, relational comprehension, and interpretive reasoning. MA-Bench contains 12,000 structured question-answer pairs, enabling systematic assessment of both recognition accuracy and action interpretation. The results of 23 representative MLLMs reveal that there are significant challenges in capturing motion granularity and fine-grained body-part dynamics. To address these challenges, we further construct MA-Bench-Train, a large-scale training corpus with 20.5K videos annotated with structured micro-action captions for fine-tuning MLLMs. The results of Qwen3-VL-8B fine-tuned on MA-Bench-Train show clear performance improvements across micro-action reasoning and explanation tasks. Our work aims to establish a foundation benchmark for advancing MLLMs in understanding subtle micro-action and human-related behaviors. Project Page: https://MA-Bench.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2603_26586
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MA-Bench: Towards Fine-grained Micro-Action Understanding
Li, Kun
Gu, Jihao
Wang, Fei
Wu, Zhiliang
Fan, Hehe
Guo, Dan
Computer Vision and Pattern Recognition
With the rapid development of Multimodal Large Language Models (MLLMs), their potential in Micro-Action understanding, a vital role in human emotion analysis, remains unexplored due to the absence of specialized benchmarks. To tackle this issue, we present MA-Bench, a benchmark comprising 1,000 videos and a three-tier evaluation architecture that progressively examines micro-action perception, relational comprehension, and interpretive reasoning. MA-Bench contains 12,000 structured question-answer pairs, enabling systematic assessment of both recognition accuracy and action interpretation. The results of 23 representative MLLMs reveal that there are significant challenges in capturing motion granularity and fine-grained body-part dynamics. To address these challenges, we further construct MA-Bench-Train, a large-scale training corpus with 20.5K videos annotated with structured micro-action captions for fine-tuning MLLMs. The results of Qwen3-VL-8B fine-tuned on MA-Bench-Train show clear performance improvements across micro-action reasoning and explanation tasks. Our work aims to establish a foundation benchmark for advancing MLLMs in understanding subtle micro-action and human-related behaviors. Project Page: https://MA-Bench.github.io
title MA-Bench: Towards Fine-grained Micro-Action Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.26586