VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Hanzhi, Sun, Boyang, Zhang, Anran, Pollefeys, Marc, Leutenegger, Stefan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GOPLA: Generalizable Object Placement Learning via Synthetic Augmentation of Human Arrangement
by: Zhong, Yao, et al.
Published: (2025)
by: Zhong, Yao, et al.
Published: (2025)
FrontierNet: Learning Visual Cues to Explore
by: Sun, Boyang, et al.
Published: (2025)
by: Sun, Boyang, et al.
Published: (2025)
Actron3D: Learning Actionable Neural Functions from Videos for Transferable Robotic Manipulation
by: Zhang, Anran, et al.
Published: (2025)
by: Zhang, Anran, et al.
Published: (2025)
REACT3D: Recovering Articulations for Interactive Physical 3D Scenes
by: Huang, Zhao, et al.
Published: (2025)
by: Huang, Zhao, et al.
Published: (2025)
FuncGrasp: Learning Object-Centric Neural Grasp Functions from Single Annotated Example Object
by: Chen, Hanzhi, et al.
Published: (2024)
by: Chen, Hanzhi, et al.
Published: (2024)
Memory Over Maps: 3D Object Localization Without Reconstruction
by: Zhou, Rui, et al.
Published: (2026)
by: Zhou, Rui, et al.
Published: (2026)
RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation
by: Kuang, Yuxuan, et al.
Published: (2024)
by: Kuang, Yuxuan, et al.
Published: (2024)
O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation
by: Tian, Tongxuan, et al.
Published: (2025)
by: Tian, Tongxuan, et al.
Published: (2025)
FetchBot: Learning Generalizable Object Fetching in Cluttered Scenes via Zero-Shot Sim2Real
by: Liu, Weiheng, et al.
Published: (2025)
by: Liu, Weiheng, et al.
Published: (2025)
MAPLE: Encoding Dexterous Robotic Manipulation Priors Learned From Egocentric Videos
by: Gavryushin, Alexey, et al.
Published: (2025)
by: Gavryushin, Alexey, et al.
Published: (2025)
Sight Over Site: Perception-Aware Reinforcement Learning for Efficient Robotic Inspection
by: Kuhlmann, Richard, et al.
Published: (2025)
by: Kuhlmann, Richard, et al.
Published: (2025)
Active Visual Localization for Multi-Agent Collaboration: A Data-Driven Approach
by: Hanlon, Matthew, et al.
Published: (2023)
by: Hanlon, Matthew, et al.
Published: (2023)
FUNCanon: Learning Pose-Aware Action Primitives via Functional Object Canonicalization for Generalizable Robotic Manipulation
by: Xu, Hongli, et al.
Published: (2025)
by: Xu, Hongli, et al.
Published: (2025)
Learning Generalizable 3D Manipulation With 10 Demonstrations
by: Ren, Yu, et al.
Published: (2024)
by: Ren, Yu, et al.
Published: (2024)
DROID-SLAM in the Wild
by: Li, Moyang, et al.
Published: (2026)
by: Li, Moyang, et al.
Published: (2026)
D$^3$Fields: Dynamic 3D Descriptor Fields for Zero-Shot Generalizable Rearrangement
by: Wang, Yixuan, et al.
Published: (2023)
by: Wang, Yixuan, et al.
Published: (2023)
ManiVID-3D: Generalizable View-Invariant Reinforcement Learning for Robotic Manipulation via Disentangled 3D Representations
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation
by: Wen, Youpeng, et al.
Published: (2024)
by: Wen, Youpeng, et al.
Published: (2024)
OpenFrontier: General Navigation with Visual-Language Grounded Frontiers
by: Padilla-Cerdio, Esteban, et al.
Published: (2026)
by: Padilla-Cerdio, Esteban, et al.
Published: (2026)
FlowBot3D: Learning 3D Articulation Flow to Manipulate Articulated Objects
by: Eisner, Ben, et al.
Published: (2022)
by: Eisner, Ben, et al.
Published: (2022)
DreamGrasp: Zero-Shot 3D Multi-Object Reconstruction from Partial-View Images for Robotic Manipulation
by: Kim, Young Hun, et al.
Published: (2025)
by: Kim, Young Hun, et al.
Published: (2025)
ForesightNav: Learning Scene Imagination for Efficient Exploration
by: Shah, Hardik, et al.
Published: (2025)
by: Shah, Hardik, et al.
Published: (2025)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
by: Shen, Yichao, et al.
Published: (2025)
by: Shen, Yichao, et al.
Published: (2025)
Volumetric Semantically Consistent 3D Panoptic Mapping
by: Miao, Yang, et al.
Published: (2023)
by: Miao, Yang, et al.
Published: (2023)
PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
by: Huang, Wenlong, et al.
Published: (2026)
by: Huang, Wenlong, et al.
Published: (2026)
ActLoc: Learning to Localize on the Move via Active Viewpoint Selection
by: Li, Jiajie, et al.
Published: (2025)
by: Li, Jiajie, et al.
Published: (2025)
Learning Where to Look: Self-supervised Viewpoint Selection for Active Localization using Geometrical Information
by: Di Giammarino, Luca, et al.
Published: (2024)
by: Di Giammarino, Luca, et al.
Published: (2024)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
by: Garcia, Ricardo, et al.
Published: (2024)
by: Garcia, Ricardo, et al.
Published: (2024)
Articulated 3D Scene Graphs for Open-World Mobile Manipulation
by: Büchner, Martin, et al.
Published: (2026)
by: Büchner, Martin, et al.
Published: (2026)
DriveVA: Video Action Models are Zero-Shot Drivers
by: Liu, Mengmeng, et al.
Published: (2026)
by: Liu, Mengmeng, et al.
Published: (2026)
SG-Tailor: Inter-Object Commonsense Relationship Reasoning for Scene Graph Manipulation
by: Shang, Haoliang, et al.
Published: (2025)
by: Shang, Haoliang, et al.
Published: (2025)
OpenSGA: Efficient 3D Scene Graph Alignment in the Open World
by: Chen, Gang, et al.
Published: (2026)
by: Chen, Gang, et al.
Published: (2026)
Generalizable Humanoid Manipulation with 3D Diffusion Policies
by: Ze, Yanjie, et al.
Published: (2024)
by: Ze, Yanjie, et al.
Published: (2024)
OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
by: Wang, Kuanning, et al.
Published: (2026)
by: Wang, Kuanning, et al.
Published: (2026)
Towards Generalizable Robotic Manipulation in Dynamic Environments
by: Fang, Heng, et al.
Published: (2026)
by: Fang, Heng, et al.
Published: (2026)
Towards Learning a Generalizable 3D Scene Representation from 2D Observations
by: Gromniak, Martin, et al.
Published: (2026)
by: Gromniak, Martin, et al.
Published: (2026)
Structural Action Transformer for 3D Dexterous Manipulation
by: Lei, Xiaohan, et al.
Published: (2026)
by: Lei, Xiaohan, et al.
Published: (2026)
3D-MVP: 3D Multiview Pretraining for Robotic Manipulation
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis
by: Fang, Yu, et al.
Published: (2025)
by: Fang, Yu, et al.
Published: (2025)
NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
Similar Items
-
GOPLA: Generalizable Object Placement Learning via Synthetic Augmentation of Human Arrangement
by: Zhong, Yao, et al.
Published: (2025) -
FrontierNet: Learning Visual Cues to Explore
by: Sun, Boyang, et al.
Published: (2025) -
Actron3D: Learning Actionable Neural Functions from Videos for Transferable Robotic Manipulation
by: Zhang, Anran, et al.
Published: (2025) -
REACT3D: Recovering Articulations for Interactive Physical 3D Scenes
by: Huang, Zhao, et al.
Published: (2025) -
FuncGrasp: Learning Object-Centric Neural Grasp Functions from Single Annotated Example Object
by: Chen, Hanzhi, et al.
Published: (2024)