Zero-shot Reconstruction of In-Scene Object Manipulation from Video
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866912783169224704 |
|---|---|
| author | Lin, Dixuan Wang, Tianyou Pan, Zhuoyang Wang, Yufu Liu, Lingjie Daniilidis, Kostas |
| author_facet | Lin, Dixuan Wang, Tianyou Pan, Zhuoyang Wang, Yufu Liu, Lingjie Daniilidis, Kostas |
| contents | We build the first system to address the problem of reconstructing in-scene object manipulation from a monocular RGB video. It is challenging due to ill-posed scene reconstruction, ambiguous hand-object depth, and the need for physically plausible interactions. Existing methods operate in hand centric coordinates and ignore the scene, hindering metric accuracy and practical use. In our method, we first use data-driven foundation models to initialize the core components, including the object mesh and poses, the scene point cloud, and the hand poses. We then apply a two-stage optimization that recovers a complete hand-object motion from grasping to interaction, which remains consistent with the scene information observed in the input video. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_19684 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Zero-shot Reconstruction of In-Scene Object Manipulation from Video Lin, Dixuan Wang, Tianyou Pan, Zhuoyang Wang, Yufu Liu, Lingjie Daniilidis, Kostas Computer Vision and Pattern Recognition Robotics We build the first system to address the problem of reconstructing in-scene object manipulation from a monocular RGB video. It is challenging due to ill-posed scene reconstruction, ambiguous hand-object depth, and the need for physically plausible interactions. Existing methods operate in hand centric coordinates and ignore the scene, hindering metric accuracy and practical use. In our method, we first use data-driven foundation models to initialize the core components, including the object mesh and poses, the scene point cloud, and the hand poses. We then apply a two-stage optimization that recovers a complete hand-object motion from grasping to interaction, which remains consistent with the scene information observed in the input video. |
| title | Zero-shot Reconstruction of In-Scene Object Manipulation from Video |
| topic | Computer Vision and Pattern Recognition Robotics |
| url | https://arxiv.org/abs/2512.19684 |