MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Fan, Dong, Xingping, Yu, Xin, Luo, Wenhan, Liu, Wei, Zhang, Kaihao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910059816026112
author Yang, Fan
Dong, Xingping
Yu, Xin
Luo, Wenhan
Liu, Wei
Zhang, Kaihao
author_facet Yang, Fan
Dong, Xingping
Yu, Xin
Luo, Wenhan
Liu, Wei
Zhang, Kaihao
contents Understanding high-resolution (HR) images remains a critical challenge for multimodal large language models (MLLMs). Recent approaches leverage vision-based retrieval-augmented generation (RAG) to retrieve query-relevant crops from HR images, improving understanding capacity of MLLMs. However, this paradigm often leads to object fragmentation, resulting in semantic bias and incomplete retrieval, while also introducing false positives from irrelevant background patches. To address these issues, we propose Multi-resolution Retrieval-Detection (MRD), a training-free framework that enhances HR image understanding from both local and global perspectives. Locally, MRD enforces cross-scale semantic consistency via multi-resolution semantic fusion to mitigate single-resolution bias and alleviate object fragmentation. Globally, it integrates open-vocabulary object detection (OVD) as localization priors within a unified framework. Extensive experiments across multiple MLLMs on HR image benchmarks demonstrate that MRD achieves state-of-the-art (SOTA) performance on both single-object and multi-object understanding tasks. Code will be available at: https://github.com/yf0412/MRD.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02906
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
Yang, Fan
Dong, Xingping
Yu, Xin
Luo, Wenhan
Liu, Wei
Zhang, Kaihao
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Understanding high-resolution (HR) images remains a critical challenge for multimodal large language models (MLLMs). Recent approaches leverage vision-based retrieval-augmented generation (RAG) to retrieve query-relevant crops from HR images, improving understanding capacity of MLLMs. However, this paradigm often leads to object fragmentation, resulting in semantic bias and incomplete retrieval, while also introducing false positives from irrelevant background patches. To address these issues, we propose Multi-resolution Retrieval-Detection (MRD), a training-free framework that enhances HR image understanding from both local and global perspectives. Locally, MRD enforces cross-scale semantic consistency via multi-resolution semantic fusion to mitigate single-resolution bias and alleviate object fragmentation. Globally, it integrates open-vocabulary object detection (OVD) as localization priors within a unified framework. Extensive experiments across multiple MLLMs on HR image benchmarks demonstrate that MRD achieves state-of-the-art (SOTA) performance on both single-object and multi-object understanding tasks. Code will be available at: https://github.com/yf0412/MRD.
title MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2512.02906