FingerCap: Fine-grained Finger-level Hand Motion Captioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Xin, Zhu, Rui, Shen, Lei, Wang, Xinyu, Zhang, Kaihao, Zhu, Tianqing, Wu, Shuchen, Miao, Chenxi, Li, Weikang, Li, Yang, Xia, Deguo, Huang, Jizhou, Yu, Xin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909915980759040
author Shen, Xin
Zhu, Rui
Shen, Lei
Wang, Xinyu
Zhang, Kaihao
Zhu, Tianqing
Wu, Shuchen
Miao, Chenxi
Li, Weikang
Li, Yang
Xia, Deguo
Huang, Jizhou
Yu, Xin
author_facet Shen, Xin
Zhu, Rui
Shen, Lei
Wang, Xinyu
Zhang, Kaihao
Zhu, Tianqing
Wu, Shuchen
Miao, Chenxi
Li, Weikang
Li, Yang
Xia, Deguo
Huang, Jizhou
Yu, Xin
contents Understanding fine-grained human hand motion is fundamental to visual perception, embodied intelligence, and multimodal communication. In this work, we propose Fine-grained Finger-level Hand Motion Captioning (FingerCap), which aims to generate textual descriptions that capture detailed finger-level semantics of hand actions. To support this task, we curate FingerCap-40K, a large-scale corpus of 40K paired hand-motion videos and captions spanning two complementary sources: concise instruction-style finger motions and diverse, naturalistic hand-object interactions. To enable effective evaluation, we employ HandJudge, a LLM-based rubric that measures finger-level correctness and motion completeness. Temporal sparsity remains a fundamental bottleneck for current Video-MLLMs, since sparse RGB sampling is insufficient to capture the subtle, high-frequency dynamics underlying fine finger motions. As a simple and compute-friendly remedy, we introduce FiGOP (Finger Group-of-Pictures), which pairs each RGB keyframe with subsequent hand keypoints until the next keyframe. A lightweight temporal encoder converts the keypoints into motion embeddings and integrates them with RGB features. FiGOP adapts the classic GOP concept to finger motion, recovering fine temporal cues without increasing RGB density. Experiments on FingerCap-40K show that strong open- and closed-source Video-MLLMs still struggle with finger-level reasoning, while our FiGOP-augmented model yield consistent gains under HandJudge and human studies.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16951
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FingerCap: Fine-grained Finger-level Hand Motion Captioning
Shen, Xin
Zhu, Rui
Shen, Lei
Wang, Xinyu
Zhang, Kaihao
Zhu, Tianqing
Wu, Shuchen
Miao, Chenxi
Li, Weikang
Li, Yang
Xia, Deguo
Huang, Jizhou
Yu, Xin
Computer Vision and Pattern Recognition
Understanding fine-grained human hand motion is fundamental to visual perception, embodied intelligence, and multimodal communication. In this work, we propose Fine-grained Finger-level Hand Motion Captioning (FingerCap), which aims to generate textual descriptions that capture detailed finger-level semantics of hand actions. To support this task, we curate FingerCap-40K, a large-scale corpus of 40K paired hand-motion videos and captions spanning two complementary sources: concise instruction-style finger motions and diverse, naturalistic hand-object interactions. To enable effective evaluation, we employ HandJudge, a LLM-based rubric that measures finger-level correctness and motion completeness. Temporal sparsity remains a fundamental bottleneck for current Video-MLLMs, since sparse RGB sampling is insufficient to capture the subtle, high-frequency dynamics underlying fine finger motions. As a simple and compute-friendly remedy, we introduce FiGOP (Finger Group-of-Pictures), which pairs each RGB keyframe with subsequent hand keypoints until the next keyframe. A lightweight temporal encoder converts the keypoints into motion embeddings and integrates them with RGB features. FiGOP adapts the classic GOP concept to finger motion, recovering fine temporal cues without increasing RGB density. Experiments on FingerCap-40K show that strong open- and closed-source Video-MLLMs still struggle with finger-level reasoning, while our FiGOP-augmented model yield consistent gains under HandJudge and human studies.
title FingerCap: Fine-grained Finger-level Hand Motion Captioning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.16951