Saved in:
Bibliographic Details
Main Authors: Zhang, Zhenliang, Wang, Yuxi, Xie, Hongzhao, Zhao, Shiyun, Liu, Mingyuan, Lu, Yujie, He, Xinyi, Cheng, Zhenku, Peng, Yujia
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2509.17425
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908666363379712
author Zhang, Zhenliang
Wang, Yuxi
Xie, Hongzhao
Zhao, Shiyun
Liu, Mingyuan
Lu, Yujie
He, Xinyi
Cheng, Zhenku
Peng, Yujia
author_facet Zhang, Zhenliang
Wang, Yuxi
Xie, Hongzhao
Zhao, Shiyun
Liu, Mingyuan
Lu, Yujie
He, Xinyi
Cheng, Zhenku
Peng, Yujia
contents A key feature differentiating artificial general intelligence (AGI) from traditional AI is that AGI can perform composite tasks that require a wide range of capabilities. Although embodied agents powered by multimodal large language models (MLLMs) offer rich perceptual and interactive capabilities, it remains largely unexplored whether they can solve composite tasks. In the current work, we designed a set of composite tasks inspired by common daily activities observed in early childhood development. Within a dynamic and simulated home environment, these tasks span three core domains: object understanding, spatial intelligence, and social activity. We evaluated 17 leading proprietary and open-source MLLMs on these tasks. The results consistently showed poor performance across all three domains, indicating a substantial gap between current capabilities and general intelligence requirements. Together, our tasks offer a preliminary framework for evaluating the general capabilities of embodied agents, marking an early but significant step toward the development of embodied MLLMs and their real-world deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17425
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Multimodal Large Language Models with Daily Composite Tasks in Home Environments
Zhang, Zhenliang
Wang, Yuxi
Xie, Hongzhao
Zhao, Shiyun
Liu, Mingyuan
Lu, Yujie
He, Xinyi
Cheng, Zhenku
Peng, Yujia
Artificial Intelligence
A key feature differentiating artificial general intelligence (AGI) from traditional AI is that AGI can perform composite tasks that require a wide range of capabilities. Although embodied agents powered by multimodal large language models (MLLMs) offer rich perceptual and interactive capabilities, it remains largely unexplored whether they can solve composite tasks. In the current work, we designed a set of composite tasks inspired by common daily activities observed in early childhood development. Within a dynamic and simulated home environment, these tasks span three core domains: object understanding, spatial intelligence, and social activity. We evaluated 17 leading proprietary and open-source MLLMs on these tasks. The results consistently showed poor performance across all three domains, indicating a substantial gap between current capabilities and general intelligence requirements. Together, our tasks offer a preliminary framework for evaluating the general capabilities of embodied agents, marking an early but significant step toward the development of embodied MLLMs and their real-world deployment.
title Evaluating Multimodal Large Language Models with Daily Composite Tasks in Home Environments
topic Artificial Intelligence
url https://arxiv.org/abs/2509.17425