EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Yi, Ge, Yuying, Ge, Yixiao, Ding, Mingyu, Li, Bohao, Wang, Rui, Xu, Ruifeng, Shan, Ying, Liu, Xihui
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909220736073728
author Chen, Yi
Ge, Yuying
Ge, Yixiao
Ding, Mingyu
Li, Bohao
Wang, Rui
Xu, Ruifeng
Shan, Ying
Liu, Xihui
author_facet Chen, Yi
Ge, Yuying
Ge, Yixiao
Ding, Mingyu
Li, Bohao
Wang, Rui
Xu, Ruifeng
Shan, Ying
Liu, Xihui
contents The pursuit of artificial general intelligence (AGI) has been accelerated by Multimodal Large Language Models (MLLMs), which exhibit superior reasoning, generalization capabilities, and proficiency in processing multimodal inputs. A crucial milestone in the evolution of AGI is the attainment of human-level planning, a fundamental ability for making informed decisions in complex environments, and solving a wide range of real-world problems. Despite the impressive advancements in MLLMs, a question remains: How far are current MLLMs from achieving human-level planning? To shed light on this question, we introduce EgoPlan-Bench, a comprehensive benchmark to evaluate the planning abilities of MLLMs in real-world scenarios from an egocentric perspective, mirroring human perception. EgoPlan-Bench emphasizes the evaluation of planning capabilities of MLLMs, featuring realistic tasks, diverse action plans, and intricate visual observations. Our rigorous evaluation of a wide range of MLLMs reveals that EgoPlan-Bench poses significant challenges, highlighting a substantial scope for improvement in MLLMs to achieve human-level task planning. To facilitate this advancement, we further present EgoPlan-IT, a specialized instruction-tuning dataset that effectively enhances model performance on EgoPlan-Bench. We have made all codes, data, and a maintained benchmark leaderboard available to advance future research.
format Preprint
id arxiv_https___arxiv_org_abs_2312_06722
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
Chen, Yi
Ge, Yuying
Ge, Yixiao
Ding, Mingyu
Li, Bohao
Wang, Rui
Xu, Ruifeng
Shan, Ying
Liu, Xihui
Computer Vision and Pattern Recognition
Computation and Language
Robotics
The pursuit of artificial general intelligence (AGI) has been accelerated by Multimodal Large Language Models (MLLMs), which exhibit superior reasoning, generalization capabilities, and proficiency in processing multimodal inputs. A crucial milestone in the evolution of AGI is the attainment of human-level planning, a fundamental ability for making informed decisions in complex environments, and solving a wide range of real-world problems. Despite the impressive advancements in MLLMs, a question remains: How far are current MLLMs from achieving human-level planning? To shed light on this question, we introduce EgoPlan-Bench, a comprehensive benchmark to evaluate the planning abilities of MLLMs in real-world scenarios from an egocentric perspective, mirroring human perception. EgoPlan-Bench emphasizes the evaluation of planning capabilities of MLLMs, featuring realistic tasks, diverse action plans, and intricate visual observations. Our rigorous evaluation of a wide range of MLLMs reveals that EgoPlan-Bench poses significant challenges, highlighting a substantial scope for improvement in MLLMs to achieve human-level task planning. To facilitate this advancement, we further present EgoPlan-IT, a specialized instruction-tuning dataset that effectively enhances model performance on EgoPlan-Bench. We have made all codes, data, and a maintained benchmark leaderboard available to advance future research.
title EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
topic Computer Vision and Pattern Recognition
Computation and Language
Robotics
url https://arxiv.org/abs/2312.06722