AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Takanami, Ryosuke, Khrapchenkov, Petr, Morikuni, Shu, Arima, Jumpei, Takaba, Yuta, Maeda, Shunsuke, Okubo, Takuya, Sano, Genki, Sekioka, Satoshi, Kadoya, Aoi, Kambara, Motonari, Nishiura, Naoya, Suzuki, Haruto, Yoshimoto, Takanori, Sakamoto, Koya, Ono, Shinnosuke, Yang, Hu, Yashima, Daichi, Horo, Aoi, Motoda, Tomohiro, Chiyoma, Kensuke, Ito, Hiroshi, Fukuda, Koki, Goto, Akihito, Morinaga, Kazumi, Ikeda, Yuya, Kawada, Riko, Yoshikawa, Masaki, Kosuge, Norio, Noguchi, Yuki, Ota, Kei, Matsushima, Tatsuya, Iwasawa, Yusuke, Matsuo, Yutaka, Ogata, Tetsuya
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909814845603840
author Takanami, Ryosuke
Khrapchenkov, Petr
Morikuni, Shu
Arima, Jumpei
Takaba, Yuta
Maeda, Shunsuke
Okubo, Takuya
Sano, Genki
Sekioka, Satoshi
Kadoya, Aoi
Kambara, Motonari
Nishiura, Naoya
Suzuki, Haruto
Yoshimoto, Takanori
Sakamoto, Koya
Ono, Shinnosuke
Yang, Hu
Yashima, Daichi
Horo, Aoi
Motoda, Tomohiro
Chiyoma, Kensuke
Ito, Hiroshi
Fukuda, Koki
Goto, Akihito
Morinaga, Kazumi
Ikeda, Yuya
Kawada, Riko
Yoshikawa, Masaki
Kosuge, Norio
Noguchi, Yuki
Ota, Kei
Matsushima, Tatsuya
Iwasawa, Yusuke
Matsuo, Yutaka
Ogata, Tetsuya
author_facet Takanami, Ryosuke
Khrapchenkov, Petr
Morikuni, Shu
Arima, Jumpei
Takaba, Yuta
Maeda, Shunsuke
Okubo, Takuya
Sano, Genki
Sekioka, Satoshi
Kadoya, Aoi
Kambara, Motonari
Nishiura, Naoya
Suzuki, Haruto
Yoshimoto, Takanori
Sakamoto, Koya
Ono, Shinnosuke
Yang, Hu
Yashima, Daichi
Horo, Aoi
Motoda, Tomohiro
Chiyoma, Kensuke
Ito, Hiroshi
Fukuda, Koki
Goto, Akihito
Morinaga, Kazumi
Ikeda, Yuya
Kawada, Riko
Yoshikawa, Masaki
Kosuge, Norio
Noguchi, Yuki
Ota, Kei
Matsushima, Tatsuya
Iwasawa, Yusuke
Matsuo, Yutaka
Ogata, Tetsuya
contents As robots transition from controlled settings to unstructured human environments, building generalist agents that can reliably follow natural language instructions remains a central challenge. Progress in robust mobile manipulation requires large-scale multimodal datasets that capture contact-rich and long-horizon tasks, yet existing resources lack synchronized force-torque sensing, hierarchical annotations, and explicit failure cases. We address this gap with the AIRoA MoMa Dataset, a large-scale real-world multimodal dataset for mobile manipulation. It includes synchronized RGB images, joint states, six-axis wrist force-torque signals, and internal robot states, together with a novel two-layer annotation schema of sub-goals and primitive actions for hierarchical learning and error analysis. The initial dataset comprises 25,469 episodes (approx. 94 hours) collected with the Human Support Robot (HSR) and is fully standardized in the LeRobot v2.1 format. By uniquely integrating mobile manipulation, contact-rich interaction, and long-horizon structure, AIRoA MoMa provides a critical benchmark for advancing the next generation of Vision-Language-Action models. The first version of our dataset is now available at https://huggingface.co/datasets/airoa-org/airoa-moma .
format Preprint
id arxiv_https___arxiv_org_abs_2509_25032
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation
Takanami, Ryosuke
Khrapchenkov, Petr
Morikuni, Shu
Arima, Jumpei
Takaba, Yuta
Maeda, Shunsuke
Okubo, Takuya
Sano, Genki
Sekioka, Satoshi
Kadoya, Aoi
Kambara, Motonari
Nishiura, Naoya
Suzuki, Haruto
Yoshimoto, Takanori
Sakamoto, Koya
Ono, Shinnosuke
Yang, Hu
Yashima, Daichi
Horo, Aoi
Motoda, Tomohiro
Chiyoma, Kensuke
Ito, Hiroshi
Fukuda, Koki
Goto, Akihito
Morinaga, Kazumi
Ikeda, Yuya
Kawada, Riko
Yoshikawa, Masaki
Kosuge, Norio
Noguchi, Yuki
Ota, Kei
Matsushima, Tatsuya
Iwasawa, Yusuke
Matsuo, Yutaka
Ogata, Tetsuya
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
As robots transition from controlled settings to unstructured human environments, building generalist agents that can reliably follow natural language instructions remains a central challenge. Progress in robust mobile manipulation requires large-scale multimodal datasets that capture contact-rich and long-horizon tasks, yet existing resources lack synchronized force-torque sensing, hierarchical annotations, and explicit failure cases. We address this gap with the AIRoA MoMa Dataset, a large-scale real-world multimodal dataset for mobile manipulation. It includes synchronized RGB images, joint states, six-axis wrist force-torque signals, and internal robot states, together with a novel two-layer annotation schema of sub-goals and primitive actions for hierarchical learning and error analysis. The initial dataset comprises 25,469 episodes (approx. 94 hours) collected with the Human Support Robot (HSR) and is fully standardized in the LeRobot v2.1 format. By uniquely integrating mobile manipulation, contact-rich interaction, and long-horizon structure, AIRoA MoMa provides a critical benchmark for advancing the next generation of Vision-Language-Action models. The first version of our dataset is now available at https://huggingface.co/datasets/airoa-org/airoa-moma .
title AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.25032