EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866918522729267200 |
|---|---|
| author | Li, Jinzhao Chen, Yinuo Piao, Dongxu Pan, Panwang Yu, Yifan Wang, Dong Yan, Honglei Yue, Liang Wang, Shaofei Chen, Yixin Huang, Siyuan Liu, Miao |
| author_facet | Li, Jinzhao Chen, Yinuo Piao, Dongxu Pan, Panwang Yu, Yifan Wang, Dong Yan, Honglei Yue, Liang Wang, Shaofei Chen, Yixin Huang, Siyuan Liu, Miao |
| contents | Humans constantly reason about 3D proximity, the relations between their body and surrounding objects, to guide perception and action in daily life. Whether multimodal large language models (MLLMs) can perform such embodied 3D reasoning remains unclear. To this end, we introduce EgoProx, a benchmark for egocentric 3D proximity reasoning. We organize our tasks along a cognitive chain, covering intention, exploration, exploitation, and chain-of-actions reasoning. We also design an agent based data engine that produces diverse and consistent QA pairs at scale. We benchmark prevailing MLLMs on EgoProx and conduct additional analyses with dataset specific and task specific instruction tuning. We observe large cross-domain gains, indicating that current MLLMs contain some spatial knowledge; however, they still struggle to effectively leverage it for spatial reasoning VQA. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_24456 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy Li, Jinzhao Chen, Yinuo Piao, Dongxu Pan, Panwang Yu, Yifan Wang, Dong Yan, Honglei Yue, Liang Wang, Shaofei Chen, Yixin Huang, Siyuan Liu, Miao Computer Vision and Pattern Recognition Humans constantly reason about 3D proximity, the relations between their body and surrounding objects, to guide perception and action in daily life. Whether multimodal large language models (MLLMs) can perform such embodied 3D reasoning remains unclear. To this end, we introduce EgoProx, a benchmark for egocentric 3D proximity reasoning. We organize our tasks along a cognitive chain, covering intention, exploration, exploitation, and chain-of-actions reasoning. We also design an agent based data engine that produces diverse and consistent QA pairs at scale. We benchmark prevailing MLLMs on EgoProx and conduct additional analyses with dataset specific and task specific instruction tuning. We observe large cross-domain gains, indicating that current MLLMs contain some spatial knowledge; however, they still struggle to effectively leverage it for spatial reasoning VQA. |
| title | EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.24456 |