Zero-shot Reconstruction of In-Scene Object Manipulation from Video
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Dixuan, Wang, Tianyou, Pan, Zhuoyang, Wang, Yufu, Liu, Lingjie, Daniilidis, Kostas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos
by: Wang, Yufu, et al.
Published: (2024)
by: Wang, Yufu, et al.
Published: (2024)
ZeroMimic: Distilling Robotic Manipulation Skills from Web Videos
by: Shi, Junyao, et al.
Published: (2025)
by: Shi, Junyao, et al.
Published: (2025)
DIMO: Diverse 3D Motion Generation for Arbitrary Objects
by: Mou, Linzhan, et al.
Published: (2025)
by: Mou, Linzhan, et al.
Published: (2025)
HDGS: Textured 2D Gaussian Splatting for Enhanced Scene Rendering
by: Song, Yunzhou, et al.
Published: (2024)
by: Song, Yunzhou, et al.
Published: (2024)
StereoDiff: Stereo-Diffusion Synergy for Video Depth Estimation
by: Li, Haodong, et al.
Published: (2025)
by: Li, Haodong, et al.
Published: (2025)
Multimodal LLM Guided Exploration and Active Mapping using Fisher Information
by: Jiang, Wen, et al.
Published: (2024)
by: Jiang, Wen, et al.
Published: (2024)
Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
by: Chen, Jiahe, et al.
Published: (2026)
by: Chen, Jiahe, et al.
Published: (2026)
Continuous-Time Human Motion Field from Events
by: Wang, Ziyun, et al.
Published: (2024)
by: Wang, Ziyun, et al.
Published: (2024)
SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object Navigation
by: Yin, Hang, et al.
Published: (2024)
by: Yin, Hang, et al.
Published: (2024)
Zero-1-to-G: Taming Pretrained 2D Diffusion Model for Direct 3D Generation
by: Meng, Xuyi, et al.
Published: (2025)
by: Meng, Xuyi, et al.
Published: (2025)
PanopticRecon: Leverage Open-vocabulary Instance Segmentation for Zero-shot Panoptic Reconstruction
by: Yu, Xuan, et al.
Published: (2024)
by: Yu, Xuan, et al.
Published: (2024)
DeformGS: Scene Flow in Highly Deformable Scenes for Deformable Object Manipulation
by: Duisterhof, Bardienus P., et al.
Published: (2023)
by: Duisterhof, Bardienus P., et al.
Published: (2023)
Dexterous Manipulation Policies from RGB Human Videos via 3D Hand-Object Trajectory Reconstruction
by: Chen, Hongyi, et al.
Published: (2026)
by: Chen, Hongyi, et al.
Published: (2026)
EqNIO: Subequivariant Neural Inertial Odometry
by: Jayanth, Royina Karegoudra, et al.
Published: (2024)
by: Jayanth, Royina Karegoudra, et al.
Published: (2024)
Zero-shot Object-Centric Instruction Following: Integrating Foundation Models with Traditional Navigation
by: Raychaudhuri, Sonia, et al.
Published: (2024)
by: Raychaudhuri, Sonia, et al.
Published: (2024)
ViTa-Zero: Zero-shot Visuotactile Object 6D Pose Estimation
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
GAMMA: Generalizable Articulation Modeling and Manipulation for Articulated Objects
by: Yu, Qiaojun, et al.
Published: (2023)
by: Yu, Qiaojun, et al.
Published: (2023)
DreamGrasp: Zero-Shot 3D Multi-Object Reconstruction from Partial-View Images for Robotic Manipulation
by: Kim, Young Hun, et al.
Published: (2025)
by: Kim, Young Hun, et al.
Published: (2025)
Fast Feature Field ($\text{F}^3$): A Predictive Representation of Events
by: Das, Richeek, et al.
Published: (2025)
by: Das, Richeek, et al.
Published: (2025)
Track Everything Everywhere Fast and Robustly
by: Song, Yunzhou, et al.
Published: (2024)
by: Song, Yunzhou, et al.
Published: (2024)
Differentiable Inverse Graphics for Zero-shot Scene Reconstruction and Robot Grasping
by: Arriaga, Octavio, et al.
Published: (2026)
by: Arriaga, Octavio, et al.
Published: (2026)
VoroNav: Voronoi-based Zero-shot Object Navigation with Large Language Model
by: Wu, Pengying, et al.
Published: (2024)
by: Wu, Pengying, et al.
Published: (2024)
Local Policies Enable Zero-shot Long-horizon Manipulation
by: Dalal, Murtaza, et al.
Published: (2024)
by: Dalal, Murtaza, et al.
Published: (2024)
FetchBot: Learning Generalizable Object Fetching in Cluttered Scenes via Zero-Shot Sim2Real
by: Liu, Weiheng, et al.
Published: (2025)
by: Liu, Weiheng, et al.
Published: (2025)
Next Best Sense: Guiding Vision and Touch with FisherRF for 3D Gaussian Splatting
by: Strong, Matthew, et al.
Published: (2024)
by: Strong, Matthew, et al.
Published: (2024)
Active Next-Best-View Optimization for Risk-Averse Path Planning
by: Khass, Amirhossein Mollaei, et al.
Published: (2025)
by: Khass, Amirhossein Mollaei, et al.
Published: (2025)
Motion-prior Contrast Maximization for Dense Continuous-Time Motion Estimation
by: Hamann, Friedhelm, et al.
Published: (2024)
by: Hamann, Friedhelm, et al.
Published: (2024)
The impact of Compositionality in Zero-shot Multi-label action recognition for Object-based tasks
by: Calabrese, Carmela, et al.
Published: (2024)
by: Calabrese, Carmela, et al.
Published: (2024)
Online 3D Scene Reconstruction Using Neural Object Priors
by: Chabal, Thomas, et al.
Published: (2025)
by: Chabal, Thomas, et al.
Published: (2025)
ETAP: Event-based Tracking of Any Point
by: Hamann, Friedhelm, et al.
Published: (2024)
by: Hamann, Friedhelm, et al.
Published: (2024)
Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction
by: Kerr, Justin, et al.
Published: (2024)
by: Kerr, Justin, et al.
Published: (2024)
ZING-3D: Zero-shot Incremental 3D Scene Graphs via Vision-Language Models
by: Saxena, Pranav, et al.
Published: (2025)
by: Saxena, Pranav, et al.
Published: (2025)
Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies
by: Qian, Jianing, et al.
Published: (2024)
by: Qian, Jianing, et al.
Published: (2024)
You Only Scan Once: A Dynamic Scene Reconstruction Pipeline for 6-DoF Robotic Grasping of Novel Objects
by: Zhou, Lei, et al.
Published: (2024)
by: Zhou, Lei, et al.
Published: (2024)
PanoSLAM: Panoptic 3D Scene Reconstruction via Gaussian SLAM
by: Chen, Runnan, et al.
Published: (2024)
by: Chen, Runnan, et al.
Published: (2024)
Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos
by: Yuan, Chengbo, et al.
Published: (2024)
by: Yuan, Chengbo, et al.
Published: (2024)
FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation
by: Shao, Dian, et al.
Published: (2026)
by: Shao, Dian, et al.
Published: (2026)
PhysHMR: Learning Humanoid Control Policies from Vision for Physically Plausible Human Motion Reconstruction
by: Feng, Qiao, et al.
Published: (2025)
by: Feng, Qiao, et al.
Published: (2025)
6D Object Pose Tracking in Internet Videos for Robotic Manipulation
by: Ponimatkin, Georgy, et al.
Published: (2025)
by: Ponimatkin, Georgy, et al.
Published: (2025)
One-shot Video Imitation via Parameterized Symbolic Abstraction Graphs
by: Wang, Jianren, et al.
Published: (2024)
by: Wang, Jianren, et al.
Published: (2024)
Similar Items
-
TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos
by: Wang, Yufu, et al.
Published: (2024) -
ZeroMimic: Distilling Robotic Manipulation Skills from Web Videos
by: Shi, Junyao, et al.
Published: (2025) -
DIMO: Diverse 3D Motion Generation for Arbitrary Objects
by: Mou, Linzhan, et al.
Published: (2025) -
HDGS: Textured 2D Gaussian Splatting for Enhanced Scene Rendering
by: Song, Yunzhou, et al.
Published: (2024) -
StereoDiff: Stereo-Diffusion Synergy for Video Depth Estimation
by: Li, Haodong, et al.
Published: (2025)