HuMoCon: Concept Discovery for Human Motion Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Qihang, Tang, Chengcheng, Tekin, Bugra, Ma, Shugao, Yang, Yanchao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909624969461760
author Fang, Qihang
Tang, Chengcheng
Tekin, Bugra
Ma, Shugao
Yang, Yanchao
author_facet Fang, Qihang
Tang, Chengcheng
Tekin, Bugra
Ma, Shugao
Yang, Yanchao
contents We present HuMoCon, a novel motion-video understanding framework designed for advanced human behavior analysis. The core of our method is a human motion concept discovery framework that efficiently trains multi-modal encoders to extract semantically meaningful and generalizable features. HuMoCon addresses key challenges in motion concept discovery for understanding and reasoning, including the lack of explicit multi-modality feature alignment and the loss of high-frequency information in masked autoencoding frameworks. Our approach integrates a feature alignment strategy that leverages video for contextual understanding and motion for fine-grained interaction modeling, further with a velocity reconstruction mechanism to enhance high-frequency feature expression and mitigate temporal over-smoothing. Comprehensive experiments on standard benchmarks demonstrate that HuMoCon enables effective motion concept discovery and significantly outperforms state-of-the-art methods in training large models for human motion understanding. We will open-source the associated code with our paper.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20920
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HuMoCon: Concept Discovery for Human Motion Understanding
Fang, Qihang
Tang, Chengcheng
Tekin, Bugra
Ma, Shugao
Yang, Yanchao
Computer Vision and Pattern Recognition
68T07
I.2.10; I.2.7
We present HuMoCon, a novel motion-video understanding framework designed for advanced human behavior analysis. The core of our method is a human motion concept discovery framework that efficiently trains multi-modal encoders to extract semantically meaningful and generalizable features. HuMoCon addresses key challenges in motion concept discovery for understanding and reasoning, including the lack of explicit multi-modality feature alignment and the loss of high-frequency information in masked autoencoding frameworks. Our approach integrates a feature alignment strategy that leverages video for contextual understanding and motion for fine-grained interaction modeling, further with a velocity reconstruction mechanism to enhance high-frequency feature expression and mitigate temporal over-smoothing. Comprehensive experiments on standard benchmarks demonstrate that HuMoCon enables effective motion concept discovery and significantly outperforms state-of-the-art methods in training large models for human motion understanding. We will open-source the associated code with our paper.
title HuMoCon: Concept Discovery for Human Motion Understanding
topic Computer Vision and Pattern Recognition
68T07
I.2.10; I.2.7
url https://arxiv.org/abs/2505.20920