Saved in:
Bibliographic Details
Main Authors: Wang, Ziyu, Liu, Chenyuan, Xiang, Yushun, Zhang, Runhao, Hao, Qingbo, Lu, Hongliang, Chen, Houyu, Feng, Zhizhong, Zheng, Kaiyue, Ye, Dehao, Zeng, Xianchao, Zhou, Xinyu, Wen, Boran, Li, Jiaxin, Zhang, Mingyu, Zheng, Kecheng, Zhu, Qian, Cheng, Ran, Li, Yong-Lu
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2601.11421
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908770900115456
author Wang, Ziyu
Liu, Chenyuan
Xiang, Yushun
Zhang, Runhao
Hao, Qingbo
Lu, Hongliang
Chen, Houyu
Feng, Zhizhong
Zheng, Kaiyue
Ye, Dehao
Zeng, Xianchao
Zhou, Xinyu
Wen, Boran
Li, Jiaxin
Zhang, Mingyu
Zheng, Kecheng
Zhu, Qian
Cheng, Ran
Li, Yong-Lu
author_facet Wang, Ziyu
Liu, Chenyuan
Xiang, Yushun
Zhang, Runhao
Hao, Qingbo
Lu, Hongliang
Chen, Houyu
Feng, Zhizhong
Zheng, Kaiyue
Ye, Dehao
Zeng, Xianchao
Zhou, Xinyu
Wen, Boran
Li, Jiaxin
Zhang, Mingyu
Zheng, Kecheng
Zhu, Qian
Cheng, Ran
Li, Yong-Lu
contents Recently, with the rapid development of robot learning and imitation learning, numerous datasets and methods have emerged. However, these datasets and their task designs often lack systematic consideration and principles. This raises important questions: Do the current datasets and task designs truly advance the capabilities of robotic agents? Do evaluations on a few common tasks accurately reflect the differentiated performance of various methods proposed by different teams and evaluated on different tasks? To address these issues, we introduce the Great March 100 (\textbf{GM-100}) as the first step towards a robot learning Olympics. GM-100 consists of 100 carefully designed tasks that cover a wide range of interactions and long-tail behaviors, aiming to provide a diverse and challenging set of tasks to comprehensively evaluate the capabilities of robotic agents and promote diversity and complexity in robot dataset task designs. These tasks are developed through systematic analysis and expansion of existing task designs, combined with insights from human-object interaction primitives and object affordances. We collect a large amount of trajectory data on different robotic platforms and evaluate several baseline models. Experimental results demonstrate that the GM-100 tasks are 1) feasible to execute and 2) sufficiently challenging to effectively differentiate the performance of current VLA models. Our data and code are available at https://rhos.ai/research/gm-100.
format Preprint
id arxiv_https___arxiv_org_abs_2601_11421
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Great March 100: 100 Detail-oriented Tasks for Evaluating Embodied AI Agents
Wang, Ziyu
Liu, Chenyuan
Xiang, Yushun
Zhang, Runhao
Hao, Qingbo
Lu, Hongliang
Chen, Houyu
Feng, Zhizhong
Zheng, Kaiyue
Ye, Dehao
Zeng, Xianchao
Zhou, Xinyu
Wen, Boran
Li, Jiaxin
Zhang, Mingyu
Zheng, Kecheng
Zhu, Qian
Cheng, Ran
Li, Yong-Lu
Robotics
Artificial Intelligence
Recently, with the rapid development of robot learning and imitation learning, numerous datasets and methods have emerged. However, these datasets and their task designs often lack systematic consideration and principles. This raises important questions: Do the current datasets and task designs truly advance the capabilities of robotic agents? Do evaluations on a few common tasks accurately reflect the differentiated performance of various methods proposed by different teams and evaluated on different tasks? To address these issues, we introduce the Great March 100 (\textbf{GM-100}) as the first step towards a robot learning Olympics. GM-100 consists of 100 carefully designed tasks that cover a wide range of interactions and long-tail behaviors, aiming to provide a diverse and challenging set of tasks to comprehensively evaluate the capabilities of robotic agents and promote diversity and complexity in robot dataset task designs. These tasks are developed through systematic analysis and expansion of existing task designs, combined with insights from human-object interaction primitives and object affordances. We collect a large amount of trajectory data on different robotic platforms and evaluate several baseline models. Experimental results demonstrate that the GM-100 tasks are 1) feasible to execute and 2) sufficiently challenging to effectively differentiate the performance of current VLA models. Our data and code are available at https://rhos.ai/research/gm-100.
title The Great March 100: 100 Detail-oriented Tasks for Evaluating Embodied AI Agents
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2601.11421