AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866909814845603840 |
|---|---|
| author | Takanami, Ryosuke Khrapchenkov, Petr Morikuni, Shu Arima, Jumpei Takaba, Yuta Maeda, Shunsuke Okubo, Takuya Sano, Genki Sekioka, Satoshi Kadoya, Aoi Kambara, Motonari Nishiura, Naoya Suzuki, Haruto Yoshimoto, Takanori Sakamoto, Koya Ono, Shinnosuke Yang, Hu Yashima, Daichi Horo, Aoi Motoda, Tomohiro Chiyoma, Kensuke Ito, Hiroshi Fukuda, Koki Goto, Akihito Morinaga, Kazumi Ikeda, Yuya Kawada, Riko Yoshikawa, Masaki Kosuge, Norio Noguchi, Yuki Ota, Kei Matsushima, Tatsuya Iwasawa, Yusuke Matsuo, Yutaka Ogata, Tetsuya |
| author_facet | Takanami, Ryosuke Khrapchenkov, Petr Morikuni, Shu Arima, Jumpei Takaba, Yuta Maeda, Shunsuke Okubo, Takuya Sano, Genki Sekioka, Satoshi Kadoya, Aoi Kambara, Motonari Nishiura, Naoya Suzuki, Haruto Yoshimoto, Takanori Sakamoto, Koya Ono, Shinnosuke Yang, Hu Yashima, Daichi Horo, Aoi Motoda, Tomohiro Chiyoma, Kensuke Ito, Hiroshi Fukuda, Koki Goto, Akihito Morinaga, Kazumi Ikeda, Yuya Kawada, Riko Yoshikawa, Masaki Kosuge, Norio Noguchi, Yuki Ota, Kei Matsushima, Tatsuya Iwasawa, Yusuke Matsuo, Yutaka Ogata, Tetsuya |
| contents | As robots transition from controlled settings to unstructured human environments, building generalist agents that can reliably follow natural language instructions remains a central challenge. Progress in robust mobile manipulation requires large-scale multimodal datasets that capture contact-rich and long-horizon tasks, yet existing resources lack synchronized force-torque sensing, hierarchical annotations, and explicit failure cases. We address this gap with the AIRoA MoMa Dataset, a large-scale real-world multimodal dataset for mobile manipulation. It includes synchronized RGB images, joint states, six-axis wrist force-torque signals, and internal robot states, together with a novel two-layer annotation schema of sub-goals and primitive actions for hierarchical learning and error analysis. The initial dataset comprises 25,469 episodes (approx. 94 hours) collected with the Human Support Robot (HSR) and is fully standardized in the LeRobot v2.1 format. By uniquely integrating mobile manipulation, contact-rich interaction, and long-horizon structure, AIRoA MoMa provides a critical benchmark for advancing the next generation of Vision-Language-Action models. The first version of our dataset is now available at https://huggingface.co/datasets/airoa-org/airoa-moma . |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_25032 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation Takanami, Ryosuke Khrapchenkov, Petr Morikuni, Shu Arima, Jumpei Takaba, Yuta Maeda, Shunsuke Okubo, Takuya Sano, Genki Sekioka, Satoshi Kadoya, Aoi Kambara, Motonari Nishiura, Naoya Suzuki, Haruto Yoshimoto, Takanori Sakamoto, Koya Ono, Shinnosuke Yang, Hu Yashima, Daichi Horo, Aoi Motoda, Tomohiro Chiyoma, Kensuke Ito, Hiroshi Fukuda, Koki Goto, Akihito Morinaga, Kazumi Ikeda, Yuya Kawada, Riko Yoshikawa, Masaki Kosuge, Norio Noguchi, Yuki Ota, Kei Matsushima, Tatsuya Iwasawa, Yusuke Matsuo, Yutaka Ogata, Tetsuya Robotics Artificial Intelligence Computer Vision and Pattern Recognition As robots transition from controlled settings to unstructured human environments, building generalist agents that can reliably follow natural language instructions remains a central challenge. Progress in robust mobile manipulation requires large-scale multimodal datasets that capture contact-rich and long-horizon tasks, yet existing resources lack synchronized force-torque sensing, hierarchical annotations, and explicit failure cases. We address this gap with the AIRoA MoMa Dataset, a large-scale real-world multimodal dataset for mobile manipulation. It includes synchronized RGB images, joint states, six-axis wrist force-torque signals, and internal robot states, together with a novel two-layer annotation schema of sub-goals and primitive actions for hierarchical learning and error analysis. The initial dataset comprises 25,469 episodes (approx. 94 hours) collected with the Human Support Robot (HSR) and is fully standardized in the LeRobot v2.1 format. By uniquely integrating mobile manipulation, contact-rich interaction, and long-horizon structure, AIRoA MoMa provides a critical benchmark for advancing the next generation of Vision-Language-Action models. The first version of our dataset is now available at https://huggingface.co/datasets/airoa-org/airoa-moma . |
| title | AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation |
| topic | Robotics Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2509.25032 |