Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zhaochen, Gao, Kaiwen, Liang, Shuyi, Xiao, Bin, Qiao, Limeng, Ma, Lin, Jiang, Tingting |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PLUG: Revisiting Amodal Segmentation with Foundation Model and Hierarchical Focus
von: Liu, Zhaochen, et al.
Veröffentlicht: (2024)
von: Liu, Zhaochen, et al.
Veröffentlicht: (2024)
GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning
von: Liu, Zhaochen, et al.
Veröffentlicht: (2026)
von: Liu, Zhaochen, et al.
Veröffentlicht: (2026)
STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
von: Qin, Jie, et al.
Veröffentlicht: (2025)
von: Qin, Jie, et al.
Veröffentlicht: (2025)
RIV: Recursive Introspection Mask Diffusion Vision Language Model
von: Li, YuQian, et al.
Veröffentlicht: (2025)
von: Li, YuQian, et al.
Veröffentlicht: (2025)
BLADE: Box-Level Supervised Amodal Segmentation through Directed Expansion
von: Liu, Zhaochen, et al.
Veröffentlicht: (2024)
von: Liu, Zhaochen, et al.
Veröffentlicht: (2024)
VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
von: Wang, Ziteng, et al.
Veröffentlicht: (2025)
von: Wang, Ziteng, et al.
Veröffentlicht: (2025)
Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
von: Zhang, Fan, et al.
Veröffentlicht: (2025)
von: Zhang, Fan, et al.
Veröffentlicht: (2025)
MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
von: Jiang, Xi, et al.
Veröffentlicht: (2024)
von: Jiang, Xi, et al.
Veröffentlicht: (2024)
Amodal Segmentation for Laparoscopic Surgery Video Instruments
von: Shi, Ruohua, et al.
Veröffentlicht: (2024)
von: Shi, Ruohua, et al.
Veröffentlicht: (2024)
AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
Learning Trajectory-Aware Multimodal Large Language Models for Video Reasoning Segmentation
von: Luo, Jingnan, et al.
Veröffentlicht: (2026)
von: Luo, Jingnan, et al.
Veröffentlicht: (2026)
VideoAesBench: Benchmarking the Video Aesthetics Perception Capabilities of Large Multimodal Models
von: Li, Yunhao, et al.
Veröffentlicht: (2026)
von: Li, Yunhao, et al.
Veröffentlicht: (2026)
Q-Bench-Portrait: Benchmarking Multimodal Large Language Models on Portrait Image Quality Perception
von: Wu, Sijing, et al.
Veröffentlicht: (2026)
von: Wu, Sijing, et al.
Veröffentlicht: (2026)
MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models
von: Liu, Xin, et al.
Veröffentlicht: (2023)
von: Liu, Xin, et al.
Veröffentlicht: (2023)
Introducing Visual Perception Token into Multimodal Large Language Model
von: Yu, Runpeng, et al.
Veröffentlicht: (2025)
von: Yu, Runpeng, et al.
Veröffentlicht: (2025)
UniComp: Rethinking Video Compression Through Informational Uniqueness
von: Yuan, Chao, et al.
Veröffentlicht: (2025)
von: Yuan, Chao, et al.
Veröffentlicht: (2025)
DesignProbe: A Graphic Design Benchmark for Multimodal Large Language Models
von: Lin, Jieru, et al.
Veröffentlicht: (2024)
von: Lin, Jieru, et al.
Veröffentlicht: (2024)
VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
von: Zhao, Fufangchen, et al.
Veröffentlicht: (2025)
von: Zhao, Fufangchen, et al.
Veröffentlicht: (2025)
Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues
von: Fu, Junchen, et al.
Veröffentlicht: (2026)
von: Fu, Junchen, et al.
Veröffentlicht: (2026)
MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2025)
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2025)
Bringing Multimodal Large Language Models to Infrared-Visible Image Fusion Quality Assessment
von: Guo, Yuchen, et al.
Veröffentlicht: (2026)
von: Guo, Yuchen, et al.
Veröffentlicht: (2026)
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
von: Ying, Kaining, et al.
Veröffentlicht: (2024)
von: Ying, Kaining, et al.
Veröffentlicht: (2024)
ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models
von: Xue, Kaiwen, et al.
Veröffentlicht: (2026)
von: Xue, Kaiwen, et al.
Veröffentlicht: (2026)
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
von: Fu, Chaoyou, et al.
Veröffentlicht: (2023)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2023)
MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning
von: Wang, Chenyu, et al.
Veröffentlicht: (2024)
von: Wang, Chenyu, et al.
Veröffentlicht: (2024)
FaceInsight: A Multimodal Large Language Model for Face Perception
von: Li, Jingzhi, et al.
Veröffentlicht: (2025)
von: Li, Jingzhi, et al.
Veröffentlicht: (2025)
Rethinking Occlusion in FER: A Semantic-Aware Perspective and Go Beyond
von: Zhai, Huiyu, et al.
Veröffentlicht: (2025)
von: Zhai, Huiyu, et al.
Veröffentlicht: (2025)
INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
Safety of Multimodal Large Language Models on Images and Texts
von: Liu, Xin, et al.
Veröffentlicht: (2024)
von: Liu, Xin, et al.
Veröffentlicht: (2024)
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models
von: De Min, Thomas, et al.
Veröffentlicht: (2026)
von: De Min, Thomas, et al.
Veröffentlicht: (2026)
RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition
von: Wang, Binhao, et al.
Veröffentlicht: (2026)
von: Wang, Binhao, et al.
Veröffentlicht: (2026)
Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection
von: Tang, Lv, et al.
Veröffentlicht: (2023)
von: Tang, Lv, et al.
Veröffentlicht: (2023)
Semantic Alignment for Multimodal Large Language Models
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
SHIELD : An Evaluation Benchmark for Face Spoofing and Forgery Detection with Multimodal Large Language Models
von: Shi, Yichen, et al.
Veröffentlicht: (2024)
von: Shi, Yichen, et al.
Veröffentlicht: (2024)
Control Color: Multimodal Diffusion-based Interactive Image Colorization
von: Liang, Zhexin, et al.
Veröffentlicht: (2024)
von: Liang, Zhexin, et al.
Veröffentlicht: (2024)
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
von: Gao, Hongcheng, et al.
Veröffentlicht: (2025)
von: Gao, Hongcheng, et al.
Veröffentlicht: (2025)
Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PLUG: Revisiting Amodal Segmentation with Foundation Model and Hierarchical Focus
von: Liu, Zhaochen, et al.
Veröffentlicht: (2024) -
GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning
von: Liu, Zhaochen, et al.
Veröffentlicht: (2026) -
STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
von: Qin, Jie, et al.
Veröffentlicht: (2025) -
RIV: Recursive Introspection Mask Diffusion Vision Language Model
von: Li, YuQian, et al.
Veröffentlicht: (2025) -
BLADE: Box-Level Supervised Amodal Segmentation through Directed Expansion
von: Liu, Zhaochen, et al.
Veröffentlicht: (2024)