Vision-based Manipulation from Single Human Video with Open-World Object Graphs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Yifeng, Lim, Arisrei, Stone, Peter, Zhu, Yuke |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LOTUS: Continual Imitation Learning for Robot Manipulation Through Unsupervised Skill Discovery
by: Wan, Weikang, et al.
Published: (2023)
by: Wan, Weikang, et al.
Published: (2023)
OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation
by: Li, Jinhan, et al.
Published: (2024)
by: Li, Jinhan, et al.
Published: (2024)
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
by: Chen, Zerui, et al.
Published: (2024)
by: Chen, Zerui, et al.
Published: (2024)
Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
by: Kawaharazuka, Kento, et al.
Published: (2025)
by: Kawaharazuka, Kento, et al.
Published: (2025)
Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning
by: Escoriza, Adrià López, et al.
Published: (2025)
by: Escoriza, Adrià López, et al.
Published: (2025)
Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids
by: Lin, Toru, et al.
Published: (2025)
by: Lin, Toru, et al.
Published: (2025)
H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
by: Bi, Hongzhe, et al.
Published: (2025)
by: Bi, Hongzhe, et al.
Published: (2025)
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
by: Cheang, Chi-Lam, et al.
Published: (2024)
by: Cheang, Chi-Lam, et al.
Published: (2024)
A Super-human Vision-based Reinforcement Learning Agent for Autonomous Racing in Gran Turismo
by: Vasco, Miguel, et al.
Published: (2024)
by: Vasco, Miguel, et al.
Published: (2024)
Adaptive Mobile Manipulation for Articulated Objects In the Open World
by: Xiong, Haoyu, et al.
Published: (2024)
by: Xiong, Haoyu, et al.
Published: (2024)
DexMan: Learning Bimanual Dexterous Manipulation from Human and Generated Videos
by: Hsieh, Jhen, et al.
Published: (2025)
by: Hsieh, Jhen, et al.
Published: (2025)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
by: Zhu, Minjie, et al.
Published: (2025)
by: Zhu, Minjie, et al.
Published: (2025)
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
by: Xu, Siyu, et al.
Published: (2025)
by: Xu, Siyu, et al.
Published: (2025)
Show, Don't Tell: Detecting Novel Objects by Watching Human Videos
by: Akl, James, et al.
Published: (2026)
by: Akl, James, et al.
Published: (2026)
PointVLA: Injecting the 3D World into Vision-Language-Action Models
by: Li, Chengmeng, et al.
Published: (2025)
by: Li, Chengmeng, et al.
Published: (2025)
Entity-Centric Reinforcement Learning for Object Manipulation from Pixels
by: Haramati, Dan, et al.
Published: (2024)
by: Haramati, Dan, et al.
Published: (2024)
Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models
by: Wu, Qi, et al.
Published: (2024)
by: Wu, Qi, et al.
Published: (2024)
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
by: Luo, Hao, et al.
Published: (2025)
by: Luo, Hao, et al.
Published: (2025)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
by: Li, Qixiu, et al.
Published: (2025)
by: Li, Qixiu, et al.
Published: (2025)
DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
by: Gao, Shenyuan, et al.
Published: (2026)
by: Gao, Shenyuan, et al.
Published: (2026)
DexMimicGen: Automated Data Generation for Bimanual Dexterous Manipulation via Imitation Learning
by: Jiang, Zhenyu, et al.
Published: (2024)
by: Jiang, Zhenyu, et al.
Published: (2024)
WorldEval: World Model as Real-World Robot Policies Evaluator
by: Li, Yaxuan, et al.
Published: (2025)
by: Li, Yaxuan, et al.
Published: (2025)
Evaluating Real-World Robot Manipulation Policies in Simulation
by: Li, Xuanlin, et al.
Published: (2024)
by: Li, Xuanlin, et al.
Published: (2024)
ZeroMimic: Distilling Robotic Manipulation Skills from Web Videos
by: Shi, Junyao, et al.
Published: (2025)
by: Shi, Junyao, et al.
Published: (2025)
Surfer: Progressive Reasoning with World Models for Robotic Manipulation
by: Ren, Pengzhen, et al.
Published: (2023)
by: Ren, Pengzhen, et al.
Published: (2023)
EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
by: Hoque, Ryan, et al.
Published: (2025)
by: Hoque, Ryan, et al.
Published: (2025)
AdaptiGraph: Material-Adaptive Graph-Based Neural Dynamics for Robotic Manipulation
by: Zhang, Kaifeng, et al.
Published: (2024)
by: Zhang, Kaifeng, et al.
Published: (2024)
Vidar: Embodied Video Diffusion Model for Generalist Manipulation
by: Feng, Yao, et al.
Published: (2025)
by: Feng, Yao, et al.
Published: (2025)
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
by: Nasiriany, Soroush, et al.
Published: (2024)
by: Nasiriany, Soroush, et al.
Published: (2024)
Learning Action-Conditional and Object-Centric Gaussian Splatting World Models for Rigid Objects
by: Kreber, Jens U., et al.
Published: (2026)
by: Kreber, Jens U., et al.
Published: (2026)
DoughNet: A Visual Predictive Model for Topological Manipulation of Deformable Objects
by: Bauer, Dominik, et al.
Published: (2024)
by: Bauer, Dominik, et al.
Published: (2024)
SCENEREPLICA: Benchmarking Real-World Robot Manipulation by Creating Replicable Scenes
by: Khargonkar, Ninad, et al.
Published: (2023)
by: Khargonkar, Ninad, et al.
Published: (2023)
SG-Tailor: Inter-Object Commonsense Relationship Reasoning for Scene Graph Manipulation
by: Shang, Haoliang, et al.
Published: (2025)
by: Shang, Haoliang, et al.
Published: (2025)
ComposableNav: Instruction-Following Navigation in Dynamic Environments via Composable Diffusion
by: Hu, Zichao, et al.
Published: (2025)
by: Hu, Zichao, et al.
Published: (2025)
iVideoGPT: Interactive VideoGPTs are Scalable World Models
by: Wu, Jialong, et al.
Published: (2024)
by: Wu, Jialong, et al.
Published: (2024)
ManiSkill-HAB: A Benchmark for Low-Level Manipulation in Home Rearrangement Tasks
by: Shukla, Arth, et al.
Published: (2024)
by: Shukla, Arth, et al.
Published: (2024)
MARVL: Multi-Stage Guidance for Robotic Manipulation via Vision-Language Models
by: Zhou, Xunlan, et al.
Published: (2026)
by: Zhou, Xunlan, et al.
Published: (2026)
Planning-Guided Diffusion Policy Learning for Generalizable Contact-Rich Bimanual Manipulation
by: Li, Xuanlin, et al.
Published: (2024)
by: Li, Xuanlin, et al.
Published: (2024)
AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
by: AgiBot-World-Contributors, et al.
Published: (2025)
by: AgiBot-World-Contributors, et al.
Published: (2025)
Learning Part-Aware Dense 3D Feature Field for Generalizable Articulated Object Manipulation
by: Chen, Yue, et al.
Published: (2026)
by: Chen, Yue, et al.
Published: (2026)
Similar Items
-
LOTUS: Continual Imitation Learning for Robot Manipulation Through Unsupervised Skill Discovery
by: Wan, Weikang, et al.
Published: (2023) -
OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation
by: Li, Jinhan, et al.
Published: (2024) -
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
by: Chen, Zerui, et al.
Published: (2024) -
Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
by: Kawaharazuka, Kento, et al.
Published: (2025) -
Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning
by: Escoriza, Adrià López, et al.
Published: (2025)