Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Zhaochen, Gao, Kaiwen, Liang, Shuyi, Xiao, Bin, Qiao, Limeng, Ma, Lin, Jiang, Tingting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915430426214400
author Liu, Zhaochen
Gao, Kaiwen
Liang, Shuyi
Xiao, Bin
Qiao, Limeng
Ma, Lin
Jiang, Tingting
author_facet Liu, Zhaochen
Gao, Kaiwen
Liang, Shuyi
Xiao, Bin
Qiao, Limeng
Ma, Lin
Jiang, Tingting
contents Occlusion perception, a critical foundation for human-level spatial understanding, embodies the challenge of integrating visual recognition and reasoning. Though multimodal large language models (MLLMs) have demonstrated remarkable capabilities, their performance on occlusion perception remains under-explored. To address this gap, we introduce O-Bench, the first visual question answering (VQA) benchmark specifically designed for occlusion perception. Based on SA-1B, we construct 1,365 images featuring semantically coherent occlusion scenarios through a novel layered synthesis approach. Upon this foundation, we annotate 4,588 question-answer pairs in total across five tailored tasks, employing a reliable, semi-automatic workflow. Our extensive evaluation of 22 representative MLLMs against the human baseline reveals a significant performance gap between current MLLMs and humans, which, we find, cannot be sufficiently bridged by model scaling or thinking process. We further identify three typical failure patterns, including an overly conservative bias, a fragile gestalt prediction, and a struggle with quantitative tasks. We believe O-Bench can not only provide a vital evaluation tool for occlusion perception, but also inspire the development of MLLMs for better visual intelligence. Our benchmark will be made publicly available upon paper publication.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04059
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
Liu, Zhaochen
Gao, Kaiwen
Liang, Shuyi
Xiao, Bin
Qiao, Limeng
Ma, Lin
Jiang, Tingting
Computer Vision and Pattern Recognition
Occlusion perception, a critical foundation for human-level spatial understanding, embodies the challenge of integrating visual recognition and reasoning. Though multimodal large language models (MLLMs) have demonstrated remarkable capabilities, their performance on occlusion perception remains under-explored. To address this gap, we introduce O-Bench, the first visual question answering (VQA) benchmark specifically designed for occlusion perception. Based on SA-1B, we construct 1,365 images featuring semantically coherent occlusion scenarios through a novel layered synthesis approach. Upon this foundation, we annotate 4,588 question-answer pairs in total across five tailored tasks, employing a reliable, semi-automatic workflow. Our extensive evaluation of 22 representative MLLMs against the human baseline reveals a significant performance gap between current MLLMs and humans, which, we find, cannot be sufficiently bridged by model scaling or thinking process. We further identify three typical failure patterns, including an overly conservative bias, a fragile gestalt prediction, and a struggle with quantitative tasks. We believe O-Bench can not only provide a vital evaluation tool for occlusion perception, but also inspire the development of MLLMs for better visual intelligence. Our benchmark will be made publicly available upon paper publication.
title Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.04059