World Models for Learning Dexterous Hand-Object Interactions from Human Videos
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912969083846656 |
|---|---|
| author | Goswami, Raktim Gautam Bar, Amir Fan, David Yang, Tsung-Yen Zhou, Gaoyue Krishnamurthy, Prashanth Rabbat, Michael Khorrami, Farshad LeCun, Yann |
| author_facet | Goswami, Raktim Gautam Bar, Amir Fan, David Yang, Tsung-Yen Zhou, Gaoyue Krishnamurthy, Prashanth Rabbat, Michael Khorrami, Farshad LeCun, Yann |
| contents | Modeling dexterous hand-object interactions is challenging as it requires understanding how subtle finger motions influence the environment through contact with objects. While recent world models address interaction modeling, they typically rely on coarse action spaces that fail to capture fine-grained dexterity. We, therefore, introduce DexWM, a Dexterous Interaction World Model that predicts future latent states of the environment conditioned on past states and dexterous actions. To overcome the scarcity of finely annotated dexterous datasets, DexWM represents actions using finger keypoints extracted from egocentric videos, enabling training on over 900 hours of human and non-dexterous robot data. Further, to accurately model dexterity, we find that predicting visual features alone is insufficient; therefore, we incorporate an auxiliary hand consistency loss that enforces accurate hand configurations. DexWM outperforms prior world models conditioned on text, navigation, or full-body actions in future-state prediction and demonstrates strong zero-shot transfer to unseen skills on a Franka Panda arm with an Allegro gripper, surpassing Diffusion Policy by over 50% on average across grasping, placing, and reaching tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_13644 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | World Models for Learning Dexterous Hand-Object Interactions from Human Videos Goswami, Raktim Gautam Bar, Amir Fan, David Yang, Tsung-Yen Zhou, Gaoyue Krishnamurthy, Prashanth Rabbat, Michael Khorrami, Farshad LeCun, Yann Robotics Artificial Intelligence Computer Vision and Pattern Recognition Modeling dexterous hand-object interactions is challenging as it requires understanding how subtle finger motions influence the environment through contact with objects. While recent world models address interaction modeling, they typically rely on coarse action spaces that fail to capture fine-grained dexterity. We, therefore, introduce DexWM, a Dexterous Interaction World Model that predicts future latent states of the environment conditioned on past states and dexterous actions. To overcome the scarcity of finely annotated dexterous datasets, DexWM represents actions using finger keypoints extracted from egocentric videos, enabling training on over 900 hours of human and non-dexterous robot data. Further, to accurately model dexterity, we find that predicting visual features alone is insufficient; therefore, we incorporate an auxiliary hand consistency loss that enforces accurate hand configurations. DexWM outperforms prior world models conditioned on text, navigation, or full-body actions in future-state prediction and demonstrates strong zero-shot transfer to unseen skills on a Franka Panda arm with an Allegro gripper, surpassing Diffusion Policy by over 50% on average across grasping, placing, and reaching tasks. |
| title | World Models for Learning Dexterous Hand-Object Interactions from Human Videos |
| topic | Robotics Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.13644 |