OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jawaid, Ahad, Xiang, Yu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912574169153536
author Jawaid, Ahad
Xiang, Yu
author_facet Jawaid, Ahad
Xiang, Yu
contents Egocentric human videos provide scalable demonstrations for imitation learning, but existing corpora often lack either fine-grained, temporally localized action descriptions or dexterous hand annotations. We introduce OpenEgo, a multimodal egocentric manipulation dataset with standardized hand-pose annotations and intention-aligned action primitives. OpenEgo totals 1107 hours across six public datasets, covering 290 manipulation tasks in 600+ environments. We unify hand-pose layouts and provide descriptive, timestamped action primitives. To validate its utility, we train language-conditioned imitation-learning policies to predict dexterous hand trajectories. OpenEgo is designed to lower the barrier to learning dexterous manipulation from egocentric video and to support reproducible research in vision-language-action learning. All resources and instructions will be released at www.openegocentric.com.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05513
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation
Jawaid, Ahad
Xiang, Yu
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Egocentric human videos provide scalable demonstrations for imitation learning, but existing corpora often lack either fine-grained, temporally localized action descriptions or dexterous hand annotations. We introduce OpenEgo, a multimodal egocentric manipulation dataset with standardized hand-pose annotations and intention-aligned action primitives. OpenEgo totals 1107 hours across six public datasets, covering 290 manipulation tasks in 600+ environments. We unify hand-pose layouts and provide descriptive, timestamped action primitives. To validate its utility, we train language-conditioned imitation-learning policies to predict dexterous hand trajectories. OpenEgo is designed to lower the barrier to learning dexterous manipulation from egocentric video and to support reproducible research in vision-language-action learning. All resources and instructions will be released at www.openegocentric.com.
title OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2509.05513