Zero-shot Reconstruction of In-Scene Object Manipulation from Video

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lin, Dixuan, Wang, Tianyou, Pan, Zhuoyang, Wang, Yufu, Liu, Lingjie, Daniilidis, Kostas
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912783169224704
author Lin, Dixuan
Wang, Tianyou
Pan, Zhuoyang
Wang, Yufu
Liu, Lingjie
Daniilidis, Kostas
author_facet Lin, Dixuan
Wang, Tianyou
Pan, Zhuoyang
Wang, Yufu
Liu, Lingjie
Daniilidis, Kostas
contents We build the first system to address the problem of reconstructing in-scene object manipulation from a monocular RGB video. It is challenging due to ill-posed scene reconstruction, ambiguous hand-object depth, and the need for physically plausible interactions. Existing methods operate in hand centric coordinates and ignore the scene, hindering metric accuracy and practical use. In our method, we first use data-driven foundation models to initialize the core components, including the object mesh and poses, the scene point cloud, and the hand poses. We then apply a two-stage optimization that recovers a complete hand-object motion from grasping to interaction, which remains consistent with the scene information observed in the input video.
format Preprint
id arxiv_https___arxiv_org_abs_2512_19684
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zero-shot Reconstruction of In-Scene Object Manipulation from Video
Lin, Dixuan
Wang, Tianyou
Pan, Zhuoyang
Wang, Yufu
Liu, Lingjie
Daniilidis, Kostas
Computer Vision and Pattern Recognition
Robotics
We build the first system to address the problem of reconstructing in-scene object manipulation from a monocular RGB video. It is challenging due to ill-posed scene reconstruction, ambiguous hand-object depth, and the need for physically plausible interactions. Existing methods operate in hand centric coordinates and ignore the scene, hindering metric accuracy and practical use. In our method, we first use data-driven foundation models to initialize the core components, including the object mesh and poses, the scene point cloud, and the hand poses. We then apply a two-stage optimization that recovers a complete hand-object motion from grasping to interaction, which remains consistent with the scene information observed in the input video.
title Zero-shot Reconstruction of In-Scene Object Manipulation from Video
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2512.19684