MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916071732150272 |
|---|---|
| author | Palanisamy, Senthil Anand, Abhishek Rathore, Satpal Singh Patnaik, Pratyush Khatana, Shubhanshu Janweja, Ekaksh |
| author_facet | Palanisamy, Senthil Anand, Abhishek Rathore, Satpal Singh Patnaik, Pratyush Khatana, Shubhanshu Janweja, Ekaksh |
| contents | Vision-language-action (VLA) models have driven demand for large-scale egocentric datasets, yet the hardware and infrastructure to collect long-horizon data remain inaccessible. Datasets today typically have episodes only a few minutes long, which fails to capture the long-horizon temporal dependencies that complex robotic task execution requires. We present MobileEgo Anywhere, a framework for collecting hour-plus egocentric trajectories on commodity mobile hardware that uses modern smartphone sensors for long-term pose tracking without the hardware barriers of traditional robotics data collection. We release three components: (1) STERA, an open-source video-processing pipeline that converts raw mobile captures into standardized, training-ready formats for VLA and foundation-model research; (2) a free mobile app that lets any user record egocentric activity; and (3) a 200-hour dataset of diverse, long-form egocentric data with persistent state tracking across 584 sessions. We further show this data is a usable training signal:mid-training a VLA on it lowers held-out action-prediction error. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_05945 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware Palanisamy, Senthil Anand, Abhishek Rathore, Satpal Singh Patnaik, Pratyush Khatana, Shubhanshu Janweja, Ekaksh Computer Vision and Pattern Recognition Computation and Language Vision-language-action (VLA) models have driven demand for large-scale egocentric datasets, yet the hardware and infrastructure to collect long-horizon data remain inaccessible. Datasets today typically have episodes only a few minutes long, which fails to capture the long-horizon temporal dependencies that complex robotic task execution requires. We present MobileEgo Anywhere, a framework for collecting hour-plus egocentric trajectories on commodity mobile hardware that uses modern smartphone sensors for long-term pose tracking without the hardware barriers of traditional robotics data collection. We release three components: (1) STERA, an open-source video-processing pipeline that converts raw mobile captures into standardized, training-ready formats for VLA and foundation-model research; (2) a free mobile app that lets any user record egocentric activity; and (3) a 200-hour dataset of diverse, long-form egocentric data with persistent state tracking across 584 sessions. We further show this data is a usable training signal:mid-training a VLA on it lowers held-out action-prediction error. |
| title | MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware |
| topic | Computer Vision and Pattern Recognition Computation and Language |
| url | https://arxiv.org/abs/2605.05945 |