Flowing from Reasoning to Motion: Learning 3D Hand Trajectory Prediction from Egocentric Human Interaction Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Mingfei, Wang, Yifan, Li, Zhengqin, Bharadhwaj, Homanga, Chen, Yujin, Qin, Chuan, Kou, Ziyi, Tian, Yuan, Whitmire, Eric, Sodhi, Rajinder, Benko, Hrvoje, Shlizerman, Eli, Liu, Yue
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917176920768512
author Chen, Mingfei
Wang, Yifan
Li, Zhengqin
Bharadhwaj, Homanga
Chen, Yujin
Qin, Chuan
Kou, Ziyi
Tian, Yuan
Whitmire, Eric
Sodhi, Rajinder
Benko, Hrvoje
Shlizerman, Eli
Liu, Yue
author_facet Chen, Mingfei
Wang, Yifan
Li, Zhengqin
Bharadhwaj, Homanga
Chen, Yujin
Qin, Chuan
Kou, Ziyi
Tian, Yuan
Whitmire, Eric
Sodhi, Rajinder
Benko, Hrvoje
Shlizerman, Eli
Liu, Yue
contents Prior works on 3D hand trajectory prediction are constrained by datasets that decouple motion from semantic supervision and by models that weakly link reasoning and action. To address these, we first present the EgoMAN dataset, a large-scale egocentric dataset for interaction stage-aware 3D hand trajectory prediction with 219K 6DoF trajectories and 3M structured QA pairs for semantic, spatial, and motion reasoning. We then introduce the EgoMAN model, a reasoning-to-motion framework that links vision-language reasoning and motion generation via a trajectory-token interface. Trained progressively to align reasoning with motion dynamics, our approach yields accurate and stage-aware trajectories with generalization across real-world scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16907
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Flowing from Reasoning to Motion: Learning 3D Hand Trajectory Prediction from Egocentric Human Interaction Videos
Chen, Mingfei
Wang, Yifan
Li, Zhengqin
Bharadhwaj, Homanga
Chen, Yujin
Qin, Chuan
Kou, Ziyi
Tian, Yuan
Whitmire, Eric
Sodhi, Rajinder
Benko, Hrvoje
Shlizerman, Eli
Liu, Yue
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Prior works on 3D hand trajectory prediction are constrained by datasets that decouple motion from semantic supervision and by models that weakly link reasoning and action. To address these, we first present the EgoMAN dataset, a large-scale egocentric dataset for interaction stage-aware 3D hand trajectory prediction with 219K 6DoF trajectories and 3M structured QA pairs for semantic, spatial, and motion reasoning. We then introduce the EgoMAN model, a reasoning-to-motion framework that links vision-language reasoning and motion generation via a trajectory-token interface. Trained progressively to align reasoning with motion dynamics, our approach yields accurate and stage-aware trajectories with generalization across real-world scenes.
title Flowing from Reasoning to Motion: Learning 3D Hand Trajectory Prediction from Egocentric Human Interaction Videos
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2512.16907