EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918507642355712 |
|---|---|
| author | Shan, Haozhe Ren, Xiancong Dong, Han Shi, Haoyuan Zhang, Yingji Hu, Jiayu Zhang, Yi Dai, Yong Shen, Bin Qu, Lizhen Xu, Zenglin Ju, Xiaozhu |
| author_facet | Shan, Haozhe Ren, Xiancong Dong, Han Shi, Haoyuan Zhang, Yingji Hu, Jiayu Zhang, Yi Dai, Yong Shen, Bin Qu, Lizhen Xu, Zenglin Ju, Xiaozhu |
| contents | While large vision-language models (VLMs) are increasingly adopted as the perceptual backbone for embodied agents, existing benchmarks often rely on question-answering or multiple-choice formats. These protocols allow models to exploit linguistic priors rather than demonstrating genuine visual grounding. To address this, we present EPIC-Bench, Embodied PerceptIon BenChmark, a fine-grained grounding benchmark designed to systematically evaluate the visual perceptual capabilities of VLMs in real-world embodied environments. Comprising 6.6k meticulously annotated tuples (Image, Text, Mask), EPIC-Bench spans 23 fine-grained tasks across three core stages of the embodied interaction pipeline: Target Localization, Navigation, and Manipulation. Extensive evaluations of over 89 leading VLMs reveal that while advanced reasoning models show promise, current VLMs universally struggle with complex visual-text alignment for physical interactions. Specifically, models exhibit critical bottlenecks in multi-target counting, part-whole relationship understanding, and affordance region detection. EPIC-Bench provides a robust foundation and actionable insights for advancing the next generation of vision-driven embodied models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_17070 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models Shan, Haozhe Ren, Xiancong Dong, Han Shi, Haoyuan Zhang, Yingji Hu, Jiayu Zhang, Yi Dai, Yong Shen, Bin Qu, Lizhen Xu, Zenglin Ju, Xiaozhu Computer Vision and Pattern Recognition While large vision-language models (VLMs) are increasingly adopted as the perceptual backbone for embodied agents, existing benchmarks often rely on question-answering or multiple-choice formats. These protocols allow models to exploit linguistic priors rather than demonstrating genuine visual grounding. To address this, we present EPIC-Bench, Embodied PerceptIon BenChmark, a fine-grained grounding benchmark designed to systematically evaluate the visual perceptual capabilities of VLMs in real-world embodied environments. Comprising 6.6k meticulously annotated tuples (Image, Text, Mask), EPIC-Bench spans 23 fine-grained tasks across three core stages of the embodied interaction pipeline: Target Localization, Navigation, and Manipulation. Extensive evaluations of over 89 leading VLMs reveal that while advanced reasoning models show promise, current VLMs universally struggle with complex visual-text alignment for physical interactions. Specifically, models exhibit critical bottlenecks in multi-target counting, part-whole relationship understanding, and affordance region detection. EPIC-Bench provides a robust foundation and actionable insights for advancing the next generation of vision-driven embodied models. |
| title | EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.17070 |