Flowing from Reasoning to Motion: Learning 3D Hand Trajectory Prediction from Egocentric Human Interaction Videos
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917176920768512 |
|---|---|
| author | Chen, Mingfei Wang, Yifan Li, Zhengqin Bharadhwaj, Homanga Chen, Yujin Qin, Chuan Kou, Ziyi Tian, Yuan Whitmire, Eric Sodhi, Rajinder Benko, Hrvoje Shlizerman, Eli Liu, Yue |
| author_facet | Chen, Mingfei Wang, Yifan Li, Zhengqin Bharadhwaj, Homanga Chen, Yujin Qin, Chuan Kou, Ziyi Tian, Yuan Whitmire, Eric Sodhi, Rajinder Benko, Hrvoje Shlizerman, Eli Liu, Yue |
| contents | Prior works on 3D hand trajectory prediction are constrained by datasets that decouple motion from semantic supervision and by models that weakly link reasoning and action. To address these, we first present the EgoMAN dataset, a large-scale egocentric dataset for interaction stage-aware 3D hand trajectory prediction with 219K 6DoF trajectories and 3M structured QA pairs for semantic, spatial, and motion reasoning. We then introduce the EgoMAN model, a reasoning-to-motion framework that links vision-language reasoning and motion generation via a trajectory-token interface. Trained progressively to align reasoning with motion dynamics, our approach yields accurate and stage-aware trajectories with generalization across real-world scenes. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_16907 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Flowing from Reasoning to Motion: Learning 3D Hand Trajectory Prediction from Egocentric Human Interaction Videos Chen, Mingfei Wang, Yifan Li, Zhengqin Bharadhwaj, Homanga Chen, Yujin Qin, Chuan Kou, Ziyi Tian, Yuan Whitmire, Eric Sodhi, Rajinder Benko, Hrvoje Shlizerman, Eli Liu, Yue Computer Vision and Pattern Recognition Artificial Intelligence Robotics Prior works on 3D hand trajectory prediction are constrained by datasets that decouple motion from semantic supervision and by models that weakly link reasoning and action. To address these, we first present the EgoMAN dataset, a large-scale egocentric dataset for interaction stage-aware 3D hand trajectory prediction with 219K 6DoF trajectories and 3M structured QA pairs for semantic, spatial, and motion reasoning. We then introduce the EgoMAN model, a reasoning-to-motion framework that links vision-language reasoning and motion generation via a trajectory-token interface. Trained progressively to align reasoning with motion dynamics, our approach yields accurate and stage-aware trajectories with generalization across real-world scenes. |
| title | Flowing from Reasoning to Motion: Learning 3D Hand Trajectory Prediction from Egocentric Human Interaction Videos |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Robotics |
| url | https://arxiv.org/abs/2512.16907 |