RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Yinglu, Lu, Zhiying, Liu, Zhihang, Sun, Yiwei, Liu, Chuanbin, Xie, Hongtao
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914211607609344
author Li, Yinglu
Lu, Zhiying
Liu, Zhihang
Sun, Yiwei
Liu, Chuanbin
Xie, Hongtao
author_facet Li, Yinglu
Lu, Zhiying
Liu, Zhihang
Sun, Yiwei
Liu, Chuanbin
Xie, Hongtao
contents Multi-modal Retrieval-Augmented Generation (RAG) has become a critical method for empowering LLMs by leveraging candidate visual documents. However, current methods consider the entire document as the basic retrieval unit, introducing substantial irrelevant visual content in two ways: 1) Relevant documents often contain large regions unrelated to the query, diluting the focus on salient information; 2) Retrieving multiple documents to increase recall further introduces redundant and irrelevant documents. These redundant contexts distract the model's attention and further degrade the performance. To address this challenge, we propose RegionRAG, a novel framework that shifts the retrieval paradigm from the document level to the region level. During training, we design a hybrid supervision strategy from both labeled data and unlabeled data to pinpoint relevant patches. During inference, we propose a dynamic pipeline that intelligently groups salient patches into complete semantic regions. By delegating the task of identifying relevant regions to the retriever, RegionRAG enables the generator to focus solely on concise, query-relevant visual content, improving both efficiency and accuracy. Experiments on six benchmarks demonstrate that RegionRAG achieves state-of-the-art performance. It improves retrieval accuracy by 10.02% in R@1 on average, and boosts question answering accuracy by 3.56% while using only 71.42% visual tokens compared with prior methods.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27261
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
Li, Yinglu
Lu, Zhiying
Liu, Zhihang
Sun, Yiwei
Liu, Chuanbin
Xie, Hongtao
Computer Vision and Pattern Recognition
Multi-modal Retrieval-Augmented Generation (RAG) has become a critical method for empowering LLMs by leveraging candidate visual documents. However, current methods consider the entire document as the basic retrieval unit, introducing substantial irrelevant visual content in two ways: 1) Relevant documents often contain large regions unrelated to the query, diluting the focus on salient information; 2) Retrieving multiple documents to increase recall further introduces redundant and irrelevant documents. These redundant contexts distract the model's attention and further degrade the performance. To address this challenge, we propose RegionRAG, a novel framework that shifts the retrieval paradigm from the document level to the region level. During training, we design a hybrid supervision strategy from both labeled data and unlabeled data to pinpoint relevant patches. During inference, we propose a dynamic pipeline that intelligently groups salient patches into complete semantic regions. By delegating the task of identifying relevant regions to the retriever, RegionRAG enables the generator to focus solely on concise, query-relevant visual content, improving both efficiency and accuracy. Experiments on six benchmarks demonstrate that RegionRAG achieves state-of-the-art performance. It improves retrieval accuracy by 10.02% in R@1 on average, and boosts question answering accuracy by 3.56% while using only 71.42% visual tokens compared with prior methods.
title RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.27261