Hand-Object Interaction Pretraining from Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Himanshu Gaurav, Loquercio, Antonio, Sferrazza, Carmelo, Wu, Jane, Qi, Haozhi, Abbeel, Pieter, Malik, Jitendra |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Twisting Lids Off with Two Hands
by: Lin, Toru, et al.
Published: (2024)
by: Lin, Toru, et al.
Published: (2024)
Deep Sensorimotor Control by Imitating Predictive Models of Human Motion
by: Singh, Himanshu Gaurav, et al.
Published: (2025)
by: Singh, Himanshu Gaurav, et al.
Published: (2025)
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
by: Goetting, Dylan, et al.
Published: (2024)
by: Goetting, Dylan, et al.
Published: (2024)
Learning Visuotactile Skills with Two Multifingered Hands
by: Lin, Toru, et al.
Published: (2024)
by: Lin, Toru, et al.
Published: (2024)
How to Peel with a Knife: Aligning Fine-Grained Manipulation with Human Preference
by: Lin, Toru, et al.
Published: (2026)
by: Lin, Toru, et al.
Published: (2026)
Object-centric 3D Motion Field for Robot Learning from Human Videos
by: Yin, Zhao-Heng, et al.
Published: (2025)
by: Yin, Zhao-Heng, et al.
Published: (2025)
DexGarmentLab: Dexterous Garment Manipulation Environment with Generalizable Policy
by: Wang, Yuran, et al.
Published: (2025)
by: Wang, Yuran, et al.
Published: (2025)
Rodrigues Network for Learning Robot Actions
by: Zhang, Jialiang, et al.
Published: (2025)
by: Zhang, Jialiang, et al.
Published: (2025)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
Reconstructing Hand-Held Objects in 3D from Images and Videos
by: Wu, Jane, et al.
Published: (2024)
by: Wu, Jane, et al.
Published: (2024)
OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
by: Huang, Huang, et al.
Published: (2025)
by: Huang, Huang, et al.
Published: (2025)
Large Video Planner Enables Generalizable Robot Control
by: Chen, Boyuan, et al.
Published: (2025)
by: Chen, Boyuan, et al.
Published: (2025)
RoSHI: A Versatile Robot-oriented Suit for Human Data In-the-Wild
by: Mao, Wenjing Margaret, et al.
Published: (2026)
by: Mao, Wenjing Margaret, et al.
Published: (2026)
SynHLMA:Synthesizing Hand Language Manipulation for Articulated Object with Discrete Human Object Interaction Representation
by: zhi, Wang, et al.
Published: (2025)
by: zhi, Wang, et al.
Published: (2025)
PhysHanDI: Physics-Based Reconstruction of Hand-Deformable Object Interactions
by: Lee, Jihyun, et al.
Published: (2026)
by: Lee, Jihyun, et al.
Published: (2026)
Closing the Visual Sim-to-Real Gap with Object-Composable NeRFs
by: Mishra, Nikhil, et al.
Published: (2024)
by: Mishra, Nikhil, et al.
Published: (2024)
The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio
by: Wang, Renhao, et al.
Published: (2025)
by: Wang, Renhao, et al.
Published: (2025)
Lightning Grasp: High Performance Procedural Grasp Synthesis with Contact Fields
by: Yin, Zhao-Heng, et al.
Published: (2025)
by: Yin, Zhao-Heng, et al.
Published: (2025)
SPIDER: Scalable Physics-Informed Dexterous Retargeting
by: Pan, Chaoyi, et al.
Published: (2025)
by: Pan, Chaoyi, et al.
Published: (2025)
HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos
by: Banerjee, Prithviraj, et al.
Published: (2024)
by: Banerjee, Prithviraj, et al.
Published: (2024)
FlowHOI: Flow-based Semantics-Grounded Generation of Hand-Object Interactions for Dexterous Robot Manipulation
by: Zeng, Huajian, et al.
Published: (2026)
by: Zeng, Huajian, et al.
Published: (2026)
Visual Imitation Enables Contextual Humanoid Control
by: Allshire, Arthur, et al.
Published: (2025)
by: Allshire, Arthur, et al.
Published: (2025)
What Matters to You? Towards Visual Representation Alignment for Robot Learning
by: Tian, Ran, et al.
Published: (2023)
by: Tian, Ran, et al.
Published: (2023)
OPENTOUCH: Bringing Full-Hand Touch to Real-World Interaction
by: Song, Yuxin Ray, et al.
Published: (2025)
by: Song, Yuxin Ray, et al.
Published: (2025)
Interactive Task Planning with Language Models
by: Li, Boyi, et al.
Published: (2023)
by: Li, Boyi, et al.
Published: (2023)
TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation
by: Zhao, Hongxiang, et al.
Published: (2025)
by: Zhao, Hongxiang, et al.
Published: (2025)
Flowing from Reasoning to Motion: Learning 3D Hand Trajectory Prediction from Egocentric Human Interaction Videos
by: Chen, Mingfei, et al.
Published: (2025)
by: Chen, Mingfei, et al.
Published: (2025)
Whole-Body Conditioned Egocentric Video Prediction
by: Bai, Yutong, et al.
Published: (2025)
by: Bai, Yutong, et al.
Published: (2025)
From Simple to Complex Skills: The Case of In-Hand Object Reorientation
by: Qi, Haozhi, et al.
Published: (2025)
by: Qi, Haozhi, et al.
Published: (2025)
Tracking by Predicting 3-D Gaussians Over Time
by: Baranwal, Tanish, et al.
Published: (2025)
by: Baranwal, Tanish, et al.
Published: (2025)
Egocentric World Model for Photorealistic Hand-Object Interaction Synthesis
by: Li, Dayou, et al.
Published: (2026)
by: Li, Dayou, et al.
Published: (2026)
Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves
by: Zhang, Xinyu, et al.
Published: (2026)
by: Zhang, Xinyu, et al.
Published: (2026)
EgoPet: Egomotion and Interaction Data from an Animal's Perspective
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
Visual Representation Learning with Stochastic Frame Prediction
by: Jang, Huiwon, et al.
Published: (2024)
by: Jang, Huiwon, et al.
Published: (2024)
Diffusion Forcing for Multi-Agent Interaction Sequence Modeling
by: Maluleke, Vongani H., et al.
Published: (2025)
by: Maluleke, Vongani H., et al.
Published: (2025)
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?
by: Majumdar, Arjun, et al.
Published: (2023)
by: Majumdar, Arjun, et al.
Published: (2023)
Maximizing Alignment with Minimal Feedback: Efficiently Learning Rewards for Visuomotor Robot Policy Alignment
by: Tian, Ran, et al.
Published: (2024)
by: Tian, Ran, et al.
Published: (2024)
D-REX: Differentiable Real-to-Sim-to-Real Engine for Learning Dexterous Grasping
by: Lou, Haozhe, et al.
Published: (2026)
by: Lou, Haozhe, et al.
Published: (2026)
MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization
by: Sukhija, Bhavya, et al.
Published: (2024)
by: Sukhija, Bhavya, et al.
Published: (2024)
Any-point Trajectory Modeling for Policy Learning
by: Wen, Chuan, et al.
Published: (2023)
by: Wen, Chuan, et al.
Published: (2023)
Similar Items
-
Twisting Lids Off with Two Hands
by: Lin, Toru, et al.
Published: (2024) -
Deep Sensorimotor Control by Imitating Predictive Models of Human Motion
by: Singh, Himanshu Gaurav, et al.
Published: (2025) -
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
by: Goetting, Dylan, et al.
Published: (2024) -
Learning Visuotactile Skills with Two Multifingered Hands
by: Lin, Toru, et al.
Published: (2024) -
How to Peel with a Knife: Aligning Fine-Grained Manipulation with Human Preference
by: Lin, Toru, et al.
Published: (2026)