EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shan, Haozhe, Ren, Xiancong, Dong, Han, Shi, Haoyuan, Zhang, Yingji, Hu, Jiayu, Zhang, Yi, Dai, Yong, Shen, Bin, Qu, Lizhen, Xu, Zenglin, Ju, Xiaozhu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918507642355712
author Shan, Haozhe
Ren, Xiancong
Dong, Han
Shi, Haoyuan
Zhang, Yingji
Hu, Jiayu
Zhang, Yi
Dai, Yong
Shen, Bin
Qu, Lizhen
Xu, Zenglin
Ju, Xiaozhu
author_facet Shan, Haozhe
Ren, Xiancong
Dong, Han
Shi, Haoyuan
Zhang, Yingji
Hu, Jiayu
Zhang, Yi
Dai, Yong
Shen, Bin
Qu, Lizhen
Xu, Zenglin
Ju, Xiaozhu
contents While large vision-language models (VLMs) are increasingly adopted as the perceptual backbone for embodied agents, existing benchmarks often rely on question-answering or multiple-choice formats. These protocols allow models to exploit linguistic priors rather than demonstrating genuine visual grounding. To address this, we present EPIC-Bench, Embodied PerceptIon BenChmark, a fine-grained grounding benchmark designed to systematically evaluate the visual perceptual capabilities of VLMs in real-world embodied environments. Comprising 6.6k meticulously annotated tuples (Image, Text, Mask), EPIC-Bench spans 23 fine-grained tasks across three core stages of the embodied interaction pipeline: Target Localization, Navigation, and Manipulation. Extensive evaluations of over 89 leading VLMs reveal that while advanced reasoning models show promise, current VLMs universally struggle with complex visual-text alignment for physical interactions. Specifically, models exhibit critical bottlenecks in multi-target counting, part-whole relationship understanding, and affordance region detection. EPIC-Bench provides a robust foundation and actionable insights for advancing the next generation of vision-driven embodied models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17070
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models
Shan, Haozhe
Ren, Xiancong
Dong, Han
Shi, Haoyuan
Zhang, Yingji
Hu, Jiayu
Zhang, Yi
Dai, Yong
Shen, Bin
Qu, Lizhen
Xu, Zenglin
Ju, Xiaozhu
Computer Vision and Pattern Recognition
While large vision-language models (VLMs) are increasingly adopted as the perceptual backbone for embodied agents, existing benchmarks often rely on question-answering or multiple-choice formats. These protocols allow models to exploit linguistic priors rather than demonstrating genuine visual grounding. To address this, we present EPIC-Bench, Embodied PerceptIon BenChmark, a fine-grained grounding benchmark designed to systematically evaluate the visual perceptual capabilities of VLMs in real-world embodied environments. Comprising 6.6k meticulously annotated tuples (Image, Text, Mask), EPIC-Bench spans 23 fine-grained tasks across three core stages of the embodied interaction pipeline: Target Localization, Navigation, and Manipulation. Extensive evaluations of over 89 leading VLMs reveal that while advanced reasoning models show promise, current VLMs universally struggle with complex visual-text alignment for physical interactions. Specifically, models exhibit critical bottlenecks in multi-target counting, part-whole relationship understanding, and affordance region detection. EPIC-Bench provides a robust foundation and actionable insights for advancing the next generation of vision-driven embodied models.
title EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.17070