Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Wenbin, Jing, Yongcheng, Ding, Liang, Wang, Yingjie, Shen, Li, Luo, Yong, Du, Bo, Tao, Dacheng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913852848865280
author Wang, Wenbin
Jing, Yongcheng
Ding, Liang
Wang, Yingjie
Shen, Li
Luo, Yong
Du, Bo
Tao, Dacheng
author_facet Wang, Wenbin
Jing, Yongcheng
Ding, Liang
Wang, Yingjie
Shen, Li
Luo, Yong
Du, Bo
Tao, Dacheng
contents High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs). To overcome the limitations of existing methods, this paper shifts away from prior dedicated heuristic approaches and revisits the most fundamental idea to HR perception by enhancing the long-context capability of MLLMs, driven by recent advances in long-context techniques like retrieval-augmented generation (RAG) for general LLMs. Towards this end, this paper presents the first study exploring the use of RAG to address HR perception challenges. Specifically, we propose Retrieval-Augmented Perception (RAP), a training-free framework that retrieves and fuses relevant image crops while preserving spatial context using the proposed Spatial-Awareness Layout. To accommodate different tasks, the proposed Retrieved-Exploration Search (RE-Search) dynamically selects the optimal number of crops based on model confidence and retrieval scores. Experimental results on HR benchmarks demonstrate the significant effectiveness of RAP, with LLaVA-v1.5-13B achieving a 43% improvement on $V^*$ Bench and 19% on HR-Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2503_01222
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG
Wang, Wenbin
Jing, Yongcheng
Ding, Liang
Wang, Yingjie
Shen, Li
Luo, Yong
Du, Bo
Tao, Dacheng
Computer Vision and Pattern Recognition
Computation and Language
High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs). To overcome the limitations of existing methods, this paper shifts away from prior dedicated heuristic approaches and revisits the most fundamental idea to HR perception by enhancing the long-context capability of MLLMs, driven by recent advances in long-context techniques like retrieval-augmented generation (RAG) for general LLMs. Towards this end, this paper presents the first study exploring the use of RAG to address HR perception challenges. Specifically, we propose Retrieval-Augmented Perception (RAP), a training-free framework that retrieves and fuses relevant image crops while preserving spatial context using the proposed Spatial-Awareness Layout. To accommodate different tasks, the proposed Retrieved-Exploration Search (RE-Search) dynamically selects the optimal number of crops based on model confidence and retrieval scores. Experimental results on HR benchmarks demonstrate the significant effectiveness of RAP, with LLaVA-v1.5-13B achieving a 43% improvement on $V^*$ Bench and 19% on HR-Bench.
title Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2503.01222