Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Bharadhwaj, Homanga, Dwibedi, Debidatta, Gupta, Abhinav, Tulsiani, Shubham, Doersch, Carl, Xiao, Ted, Shah, Dhruv, Xia, Fei, Sadigh, Dorsa, Kirmani, Sean |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
by: Bharadhwaj, Homanga, et al.
Published: (2024)
by: Bharadhwaj, Homanga, et al.
Published: (2024)
DemoDiffusion: One-Shot Human Imitation using pre-trained Diffusion Policy
by: Park, Sungjae, et al.
Published: (2025)
by: Park, Sungjae, et al.
Published: (2025)
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
by: Nasiriany, Soroush, et al.
Published: (2024)
by: Nasiriany, Soroush, et al.
Published: (2024)
RT-H: Action Hierarchies Using Language
by: Belkhale, Suneel, et al.
Published: (2024)
by: Belkhale, Suneel, et al.
Published: (2024)
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
by: Bao, Chen, et al.
Published: (2024)
by: Bao, Chen, et al.
Published: (2024)
A Taxonomy for Evaluating Generalist Robot Manipulation Policies
by: Gao, Jensen, et al.
Published: (2025)
by: Gao, Jensen, et al.
Published: (2025)
Grounding Robot Generalization in Training Data via Retrieval-Augmented VLMs
by: Gao, Jensen, et al.
Published: (2026)
by: Gao, Jensen, et al.
Published: (2026)
Semantically Controllable Augmentations for Generalizable Robot Learning
by: Chen, Zoey, et al.
Published: (2024)
by: Chen, Zoey, et al.
Published: (2024)
STEER: Flexible Robotic Manipulation via Dense Language Grounding
by: Smith, Laura, et al.
Published: (2024)
by: Smith, Laura, et al.
Published: (2024)
Efficient Data Collection for Robotic Manipulation via Compositional Generalization
by: Gao, Jensen, et al.
Published: (2024)
by: Gao, Jensen, et al.
Published: (2024)
Web2Grasp: Learning Functional Grasps from Web Images of Hand-Object Interactions
by: Chen, Hongyi, et al.
Published: (2025)
by: Chen, Hongyi, et al.
Published: (2025)
G-HOP: Generative Hand-Object Prior for Interaction Reconstruction and Grasp Synthesis
by: Ye, Yufei, et al.
Published: (2024)
by: Ye, Yufei, et al.
Published: (2024)
BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames
by: Mark, Max Sobol, et al.
Published: (2026)
by: Mark, Max Sobol, et al.
Published: (2026)
Pushing the Limits of Cross-Embodiment Learning for Manipulation and Navigation
by: Yang, Jonathan, et al.
Published: (2024)
by: Yang, Jonathan, et al.
Published: (2024)
Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization
by: Yang, Jonathan, et al.
Published: (2025)
by: Yang, Jonathan, et al.
Published: (2025)
Physically Grounded Vision-Language Models for Robotic Manipulation
by: Gao, Jensen, et al.
Published: (2023)
by: Gao, Jensen, et al.
Published: (2023)
Continual Model-Based Reinforcement Learning with Hypernetworks
by: Huang, Yizhou, et al.
Published: (2020)
by: Huang, Yizhou, et al.
Published: (2020)
ObjectForesight: Predicting Future 3D Object Trajectories from Human Videos
by: Soraki, Rustin, et al.
Published: (2026)
by: Soraki, Rustin, et al.
Published: (2026)
Distilling and Retrieving Generalizable Knowledge for Robot Manipulation via Language Corrections
by: Zha, Lihan, et al.
Published: (2023)
by: Zha, Lihan, et al.
Published: (2023)
OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
LiLo-VLA: Compositional Long-Horizon Manipulation via Linked Object-Centric Policies
by: Yang, Yue, et al.
Published: (2026)
by: Yang, Yue, et al.
Published: (2026)
FlexCap: Describe Anything in Images in Controllable Detail
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
Sparse-view Pose Estimation and Reconstruction via Analysis by Generative Synthesis
by: Zhao, Qitao, et al.
Published: (2024)
by: Zhao, Qitao, et al.
Published: (2024)
Dex4D: Task-Agnostic Point Track Policy for Sim-to-Real Dexterous Manipulation
by: Kuang, Yuxuan, et al.
Published: (2026)
by: Kuang, Yuxuan, et al.
Published: (2026)
Vision Language Models are In-Context Value Learners
by: Ma, Yecheng Jason, et al.
Published: (2024)
by: Ma, Yecheng Jason, et al.
Published: (2024)
HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies
by: Xie, Amber, et al.
Published: (2026)
by: Xie, Amber, et al.
Published: (2026)
Batch Active Learning of Reward Functions from Human Preferences
by: Bıyık, Erdem, et al.
Published: (2024)
by: Bıyık, Erdem, et al.
Published: (2024)
How to Train Your Robots? The Impact of Demonstration Modality on Imitation Learning
by: Li, Haozhuo, et al.
Published: (2025)
by: Li, Haozhuo, et al.
Published: (2025)
Imitation Bootstrapped Reinforcement Learning
by: Hu, Hengyuan, et al.
Published: (2023)
by: Hu, Hengyuan, et al.
Published: (2023)
Invariance Co-training for Robot Visual Generalization
by: Yang, Jonathan, et al.
Published: (2025)
by: Yang, Jonathan, et al.
Published: (2025)
Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning
by: Clark, Jaden, et al.
Published: (2025)
by: Clark, Jaden, et al.
Published: (2025)
Data Analogies Enable Efficient Cross-Embodiment Transfer
by: Yang, Jonathan, et al.
Published: (2026)
by: Yang, Jonathan, et al.
Published: (2026)
RT-Sketch: Goal-Conditioned Imitation Learning from Hand-Drawn Sketches
by: Sundaresan, Priya, et al.
Published: (2024)
by: Sundaresan, Priya, et al.
Published: (2024)
Evaluating Gemini Robotics Policies in a Veo World Simulator
by: Gemini Robotics Team, et al.
Published: (2025)
by: Gemini Robotics Team, et al.
Published: (2025)
Walk through Paintings: Egocentric World Models from Internet Priors
by: Bagchi, Anurag, et al.
Published: (2026)
by: Bagchi, Anurag, et al.
Published: (2026)
Dexterous Manipulation Policies from RGB Human Videos via 3D Hand-Object Trajectory Reconstruction
by: Chen, Hongyi, et al.
Published: (2026)
by: Chen, Hongyi, et al.
Published: (2026)
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
by: Ahn, Michael, et al.
Published: (2024)
by: Ahn, Michael, et al.
Published: (2024)
Action-Free Reasoning for Policy Generalization
by: Clark, Jaden, et al.
Published: (2025)
by: Clark, Jaden, et al.
Published: (2025)
Similar Items
-
Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
by: Bharadhwaj, Homanga, et al.
Published: (2024) -
DemoDiffusion: One-Shot Human Imitation using pre-trained Diffusion Policy
by: Park, Sungjae, et al.
Published: (2025) -
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
by: Nasiriany, Soroush, et al.
Published: (2024) -
RT-H: Action Hierarchies Using Language
by: Belkhale, Suneel, et al.
Published: (2024) -
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
by: Bao, Chen, et al.
Published: (2024)