FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Delitzas, Alexandros, Zhang, Chenyangguang, Gavryushin, Alexey, Di Mario, Tommaso, Sun, Boyang, Dabral, Rishabh, Guibas, Leonidas, Theobalt, Christian, Pollefeys, Marc, Engelmann, Francis, Barath, Daniel
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908992941326336
author Delitzas, Alexandros
Zhang, Chenyangguang
Gavryushin, Alexey
Di Mario, Tommaso
Sun, Boyang
Dabral, Rishabh
Guibas, Leonidas
Theobalt, Christian
Pollefeys, Marc
Engelmann, Francis
Barath, Daniel
author_facet Delitzas, Alexandros
Zhang, Chenyangguang
Gavryushin, Alexey
Di Mario, Tommaso
Sun, Boyang
Dabral, Rishabh
Guibas, Leonidas
Theobalt, Christian
Pollefeys, Marc
Engelmann, Francis
Barath, Daniel
contents We present FunRec, a method for reconstructing functional 3D digital twins of indoor scenes directly from egocentric RGB-D interaction videos. Unlike existing methods on articulated reconstruction, which rely on controlled setups, multi-state captures, or CAD priors, FunRec operates directly on in-the-wild human interaction sequences to recover interactable 3D scenes. It automatically discovers articulated parts, estimates their kinematic parameters, tracks their 3D motion, and reconstructs static and moving geometry in canonical space, yielding simulation-compatible meshes. Across new real and simulated benchmarks, FunRec surpasses prior work by a large margin, achieving up to +50 mIoU improvement in part segmentation, 5-10 times lower articulation and pose errors, and significantly higher reconstruction accuracy. We further demonstrate applications on URDF/USD export for simulation, hand-guided affordance mapping and robot-scene interaction.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05621
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos
Delitzas, Alexandros
Zhang, Chenyangguang
Gavryushin, Alexey
Di Mario, Tommaso
Sun, Boyang
Dabral, Rishabh
Guibas, Leonidas
Theobalt, Christian
Pollefeys, Marc
Engelmann, Francis
Barath, Daniel
Computer Vision and Pattern Recognition
We present FunRec, a method for reconstructing functional 3D digital twins of indoor scenes directly from egocentric RGB-D interaction videos. Unlike existing methods on articulated reconstruction, which rely on controlled setups, multi-state captures, or CAD priors, FunRec operates directly on in-the-wild human interaction sequences to recover interactable 3D scenes. It automatically discovers articulated parts, estimates their kinematic parameters, tracks their 3D motion, and reconstructs static and moving geometry in canonical space, yielding simulation-compatible meshes. Across new real and simulated benchmarks, FunRec surpasses prior work by a large margin, achieving up to +50 mIoU improvement in part segmentation, 5-10 times lower articulation and pose errors, and significantly higher reconstruction accuracy. We further demonstrate applications on URDF/USD export for simulation, hand-guided affordance mapping and robot-scene interaction.
title FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.05621