World Models for Learning Dexterous Hand-Object Interactions from Human Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Goswami, Raktim Gautam, Bar, Amir, Fan, David, Yang, Tsung-Yen, Zhou, Gaoyue, Krishnamurthy, Prashanth, Rabbat, Michael, Khorrami, Farshad, LeCun, Yann
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912969083846656
author Goswami, Raktim Gautam
Bar, Amir
Fan, David
Yang, Tsung-Yen
Zhou, Gaoyue
Krishnamurthy, Prashanth
Rabbat, Michael
Khorrami, Farshad
LeCun, Yann
author_facet Goswami, Raktim Gautam
Bar, Amir
Fan, David
Yang, Tsung-Yen
Zhou, Gaoyue
Krishnamurthy, Prashanth
Rabbat, Michael
Khorrami, Farshad
LeCun, Yann
contents Modeling dexterous hand-object interactions is challenging as it requires understanding how subtle finger motions influence the environment through contact with objects. While recent world models address interaction modeling, they typically rely on coarse action spaces that fail to capture fine-grained dexterity. We, therefore, introduce DexWM, a Dexterous Interaction World Model that predicts future latent states of the environment conditioned on past states and dexterous actions. To overcome the scarcity of finely annotated dexterous datasets, DexWM represents actions using finger keypoints extracted from egocentric videos, enabling training on over 900 hours of human and non-dexterous robot data. Further, to accurately model dexterity, we find that predicting visual features alone is insufficient; therefore, we incorporate an auxiliary hand consistency loss that enforces accurate hand configurations. DexWM outperforms prior world models conditioned on text, navigation, or full-body actions in future-state prediction and demonstrates strong zero-shot transfer to unseen skills on a Franka Panda arm with an Allegro gripper, surpassing Diffusion Policy by over 50% on average across grasping, placing, and reaching tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13644
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle World Models for Learning Dexterous Hand-Object Interactions from Human Videos
Goswami, Raktim Gautam
Bar, Amir
Fan, David
Yang, Tsung-Yen
Zhou, Gaoyue
Krishnamurthy, Prashanth
Rabbat, Michael
Khorrami, Farshad
LeCun, Yann
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Modeling dexterous hand-object interactions is challenging as it requires understanding how subtle finger motions influence the environment through contact with objects. While recent world models address interaction modeling, they typically rely on coarse action spaces that fail to capture fine-grained dexterity. We, therefore, introduce DexWM, a Dexterous Interaction World Model that predicts future latent states of the environment conditioned on past states and dexterous actions. To overcome the scarcity of finely annotated dexterous datasets, DexWM represents actions using finger keypoints extracted from egocentric videos, enabling training on over 900 hours of human and non-dexterous robot data. Further, to accurately model dexterity, we find that predicting visual features alone is insufficient; therefore, we incorporate an auxiliary hand consistency loss that enforces accurate hand configurations. DexWM outperforms prior world models conditioned on text, navigation, or full-body actions in future-state prediction and demonstrates strong zero-shot transfer to unseen skills on a Franka Panda arm with an Allegro gripper, surpassing Diffusion Policy by over 50% on average across grasping, placing, and reaching tasks.
title World Models for Learning Dexterous Hand-Object Interactions from Human Videos
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.13644