MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Guo, Shengyu, Ye, Tongrui, Zhang, Jianbo, Zhang, Zicheng, Li, Chunyi, Zhai, Guangtao
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917429000536064
author Guo, Shengyu
Ye, Tongrui
Zhang, Jianbo
Zhang, Zicheng
Li, Chunyi
Zhai, Guangtao
author_facet Guo, Shengyu
Ye, Tongrui
Zhang, Jianbo
Zhang, Zicheng
Li, Chunyi
Zhai, Guangtao
contents Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated remarkable advances in perception and reasoning, suggesting their potential for embodied intelligence. While recent studies have evaluated embodied MLLMs in interactive settings, current benchmarks mainly target capabilities to perceive, understand, and interact with external objects, lacking a systematic evaluation of self-centric intelligence. To address this, we introduce MirrorBench, a simulation-based benchmark inspired by the classical Mirror Self-Recognition (MSR) test in psychology. MirrorBench extends this paradigm to embodied MLLMs through a tiered framework of progressively challenging tasks, assessing agents from basic visual perception to high-level self-representation. Experiments on leading MLLMs show that even at the lowest level, their performance remains substantially inferior to human performance, revealing fundamental limitations in self-referential understanding. Our study bridges psychological paradigms and embodied intelligence, offering a principled framework for evaluating the emergence of general intelligence in large models. Project page: https://fflahm.github.io/mirror-bench-page/.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14785
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror
Guo, Shengyu
Ye, Tongrui
Zhang, Jianbo
Zhang, Zicheng
Li, Chunyi
Zhai, Guangtao
Artificial Intelligence
Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated remarkable advances in perception and reasoning, suggesting their potential for embodied intelligence. While recent studies have evaluated embodied MLLMs in interactive settings, current benchmarks mainly target capabilities to perceive, understand, and interact with external objects, lacking a systematic evaluation of self-centric intelligence. To address this, we introduce MirrorBench, a simulation-based benchmark inspired by the classical Mirror Self-Recognition (MSR) test in psychology. MirrorBench extends this paradigm to embodied MLLMs through a tiered framework of progressively challenging tasks, assessing agents from basic visual perception to high-level self-representation. Experiments on leading MLLMs show that even at the lowest level, their performance remains substantially inferior to human performance, revealing fundamental limitations in self-referential understanding. Our study bridges psychological paradigms and embodied intelligence, offering a principled framework for evaluating the emergence of general intelligence in large models. Project page: https://fflahm.github.io/mirror-bench-page/.
title MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror
topic Artificial Intelligence
url https://arxiv.org/abs/2604.14785