HuMoCon: Concept Discovery for Human Motion Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909624969461760 |
|---|---|
| author | Fang, Qihang Tang, Chengcheng Tekin, Bugra Ma, Shugao Yang, Yanchao |
| author_facet | Fang, Qihang Tang, Chengcheng Tekin, Bugra Ma, Shugao Yang, Yanchao |
| contents | We present HuMoCon, a novel motion-video understanding framework designed for advanced human behavior analysis. The core of our method is a human motion concept discovery framework that efficiently trains multi-modal encoders to extract semantically meaningful and generalizable features. HuMoCon addresses key challenges in motion concept discovery for understanding and reasoning, including the lack of explicit multi-modality feature alignment and the loss of high-frequency information in masked autoencoding frameworks. Our approach integrates a feature alignment strategy that leverages video for contextual understanding and motion for fine-grained interaction modeling, further with a velocity reconstruction mechanism to enhance high-frequency feature expression and mitigate temporal over-smoothing. Comprehensive experiments on standard benchmarks demonstrate that HuMoCon enables effective motion concept discovery and significantly outperforms state-of-the-art methods in training large models for human motion understanding. We will open-source the associated code with our paper. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_20920 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | HuMoCon: Concept Discovery for Human Motion Understanding Fang, Qihang Tang, Chengcheng Tekin, Bugra Ma, Shugao Yang, Yanchao Computer Vision and Pattern Recognition 68T07 I.2.10; I.2.7 We present HuMoCon, a novel motion-video understanding framework designed for advanced human behavior analysis. The core of our method is a human motion concept discovery framework that efficiently trains multi-modal encoders to extract semantically meaningful and generalizable features. HuMoCon addresses key challenges in motion concept discovery for understanding and reasoning, including the lack of explicit multi-modality feature alignment and the loss of high-frequency information in masked autoencoding frameworks. Our approach integrates a feature alignment strategy that leverages video for contextual understanding and motion for fine-grained interaction modeling, further with a velocity reconstruction mechanism to enhance high-frequency feature expression and mitigate temporal over-smoothing. Comprehensive experiments on standard benchmarks demonstrate that HuMoCon enables effective motion concept discovery and significantly outperforms state-of-the-art methods in training large models for human motion understanding. We will open-source the associated code with our paper. |
| title | HuMoCon: Concept Discovery for Human Motion Understanding |
| topic | Computer Vision and Pattern Recognition 68T07 I.2.10; I.2.7 |
| url | https://arxiv.org/abs/2505.20920 |