SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jian, Zhang, Ruiyi, Zhou, Yufan, Yu, Tong, Dernoncourt, Franck, Gu, Jiuxiang, Rossi, Ryan A., Chen, Changyou, Sun, Tong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917941943992320
author Chen, Jian
Zhang, Ruiyi
Zhou, Yufan
Yu, Tong
Dernoncourt, Franck
Gu, Jiuxiang
Rossi, Ryan A.
Chen, Changyou
Sun, Tong
author_facet Chen, Jian
Zhang, Ruiyi
Zhou, Yufan
Yu, Tong
Dernoncourt, Franck
Gu, Jiuxiang
Rossi, Ryan A.
Chen, Changyou
Sun, Tong
contents Multimodal large language models (MLLMs) have recently shown great progress in text-rich image understanding, yet they still struggle with complex, multi-page visually-rich documents. Traditional methods using document parsers for retrieval-augmented generation suffer from performance and efficiency limitations, while directly presenting all pages to MLLMs leads to inefficiencies, especially with lengthy ones. In this work, we present a novel framework named **S**elf-**V**isual **R**etrieval-**A**ugmented **G**eneration (SV-RAG), which can broaden horizons of any MLLM to support long-document understanding. We demonstrate that **MLLMs themselves can be an effective multimodal retriever** to fetch relevant pages and then answer user questions based on these pages. SV-RAG is implemented with two specific MLLM adapters, one for evidence page retrieval and the other for question answering. Empirical results show state-of-the-art performance on public benchmarks, demonstrating the effectiveness of SV-RAG.
format Preprint
id arxiv_https___arxiv_org_abs_2411_01106
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding
Chen, Jian
Zhang, Ruiyi
Zhou, Yufan
Yu, Tong
Dernoncourt, Franck
Gu, Jiuxiang
Rossi, Ryan A.
Chen, Changyou
Sun, Tong
Computer Vision and Pattern Recognition
Multimodal large language models (MLLMs) have recently shown great progress in text-rich image understanding, yet they still struggle with complex, multi-page visually-rich documents. Traditional methods using document parsers for retrieval-augmented generation suffer from performance and efficiency limitations, while directly presenting all pages to MLLMs leads to inefficiencies, especially with lengthy ones. In this work, we present a novel framework named **S**elf-**V**isual **R**etrieval-**A**ugmented **G**eneration (SV-RAG), which can broaden horizons of any MLLM to support long-document understanding. We demonstrate that **MLLMs themselves can be an effective multimodal retriever** to fetch relevant pages and then answer user questions based on these pages. SV-RAG is implemented with two specific MLLM adapters, one for evidence page retrieval and the other for question answering. Empirical results show state-of-the-art performance on public benchmarks, demonstrating the effectiveness of SV-RAG.
title SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.01106