Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Beomchan, Kim, Seongho, Kim, Hyunjun, Park, Sungjune, Ro, Yong Man
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914514917654528
author Park, Beomchan
Kim, Seongho
Kim, Hyunjun
Park, Sungjune
Ro, Yong Man
author_facet Park, Beomchan
Kim, Seongho
Kim, Hyunjun
Park, Sungjune
Ro, Yong Man
contents While Multimodal Large Language Models (MLLMs) have enhanced grounding capabilities in general scenes, their robustness in crowded scenes remains underexplored. Crowded scenes entail visual challenges (i.e., occlusion and small objects), which impair object semantics and degrade grounding performance. In contrast, language expressions are immune to such degradation and preserve object semantics. In light of these observations, we propose a novel method that overcomes such constraints by leveraging Language-Guided Semantic Cues (LGSCs). Specifically, our approach introduces a Semantic Cue Extractor (SCE) to derive semantic cues of objects from the visual pipeline of an MLLM. We then guide these cues using corresponding text embeddings to produce LGSCs as linguistic semantic priors. Subsequently, they are reintegrated into the original visual pipeline to refine object semantics. Extensive experiments and analyses demonstrate that incorporating LGSCs into an MLLM effectively improves grounding accuracy in crowded scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2604_24036
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues
Park, Beomchan
Kim, Seongho
Kim, Hyunjun
Park, Sungjune
Ro, Yong Man
Computer Vision and Pattern Recognition
Image and Video Processing
While Multimodal Large Language Models (MLLMs) have enhanced grounding capabilities in general scenes, their robustness in crowded scenes remains underexplored. Crowded scenes entail visual challenges (i.e., occlusion and small objects), which impair object semantics and degrade grounding performance. In contrast, language expressions are immune to such degradation and preserve object semantics. In light of these observations, we propose a novel method that overcomes such constraints by leveraging Language-Guided Semantic Cues (LGSCs). Specifically, our approach introduces a Semantic Cue Extractor (SCE) to derive semantic cues of objects from the visual pipeline of an MLLM. We then guide these cues using corresponding text embeddings to produce LGSCs as linguistic semantic priors. Subsequently, they are reintegrated into the original visual pipeline to refine object semantics. Extensive experiments and analyses demonstrate that incorporating LGSCs into an MLLM effectively improves grounding accuracy in crowded scenes.
title Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2604.24036