EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Jinzhao, Chen, Yinuo, Piao, Dongxu, Pan, Panwang, Yu, Yifan, Wang, Dong, Yan, Honglei, Yue, Liang, Wang, Shaofei, Chen, Yixin, Huang, Siyuan, Liu, Miao
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918522729267200
author Li, Jinzhao
Chen, Yinuo
Piao, Dongxu
Pan, Panwang
Yu, Yifan
Wang, Dong
Yan, Honglei
Yue, Liang
Wang, Shaofei
Chen, Yixin
Huang, Siyuan
Liu, Miao
author_facet Li, Jinzhao
Chen, Yinuo
Piao, Dongxu
Pan, Panwang
Yu, Yifan
Wang, Dong
Yan, Honglei
Yue, Liang
Wang, Shaofei
Chen, Yixin
Huang, Siyuan
Liu, Miao
contents Humans constantly reason about 3D proximity, the relations between their body and surrounding objects, to guide perception and action in daily life. Whether multimodal large language models (MLLMs) can perform such embodied 3D reasoning remains unclear. To this end, we introduce EgoProx, a benchmark for egocentric 3D proximity reasoning. We organize our tasks along a cognitive chain, covering intention, exploration, exploitation, and chain-of-actions reasoning. We also design an agent based data engine that produces diverse and consistent QA pairs at scale. We benchmark prevailing MLLMs on EgoProx and conduct additional analyses with dataset specific and task specific instruction tuning. We observe large cross-domain gains, indicating that current MLLMs contain some spatial knowledge; however, they still struggle to effectively leverage it for spatial reasoning VQA.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24456
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy
Li, Jinzhao
Chen, Yinuo
Piao, Dongxu
Pan, Panwang
Yu, Yifan
Wang, Dong
Yan, Honglei
Yue, Liang
Wang, Shaofei
Chen, Yixin
Huang, Siyuan
Liu, Miao
Computer Vision and Pattern Recognition
Humans constantly reason about 3D proximity, the relations between their body and surrounding objects, to guide perception and action in daily life. Whether multimodal large language models (MLLMs) can perform such embodied 3D reasoning remains unclear. To this end, we introduce EgoProx, a benchmark for egocentric 3D proximity reasoning. We organize our tasks along a cognitive chain, covering intention, exploration, exploitation, and chain-of-actions reasoning. We also design an agent based data engine that produces diverse and consistent QA pairs at scale. We benchmark prevailing MLLMs on EgoProx and conduct additional analyses with dataset specific and task specific instruction tuning. We observe large cross-domain gains, indicating that current MLLMs contain some spatial knowledge; however, they still struggle to effectively leverage it for spatial reasoning VQA.
title EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.24456