SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917941943992320 |
|---|---|
| author | Chen, Jian Zhang, Ruiyi Zhou, Yufan Yu, Tong Dernoncourt, Franck Gu, Jiuxiang Rossi, Ryan A. Chen, Changyou Sun, Tong |
| author_facet | Chen, Jian Zhang, Ruiyi Zhou, Yufan Yu, Tong Dernoncourt, Franck Gu, Jiuxiang Rossi, Ryan A. Chen, Changyou Sun, Tong |
| contents | Multimodal large language models (MLLMs) have recently shown great progress in text-rich image understanding, yet they still struggle with complex, multi-page visually-rich documents. Traditional methods using document parsers for retrieval-augmented generation suffer from performance and efficiency limitations, while directly presenting all pages to MLLMs leads to inefficiencies, especially with lengthy ones. In this work, we present a novel framework named **S**elf-**V**isual **R**etrieval-**A**ugmented **G**eneration (SV-RAG), which can broaden horizons of any MLLM to support long-document understanding. We demonstrate that **MLLMs themselves can be an effective multimodal retriever** to fetch relevant pages and then answer user questions based on these pages. SV-RAG is implemented with two specific MLLM adapters, one for evidence page retrieval and the other for question answering. Empirical results show state-of-the-art performance on public benchmarks, demonstrating the effectiveness of SV-RAG. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_01106 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding Chen, Jian Zhang, Ruiyi Zhou, Yufan Yu, Tong Dernoncourt, Franck Gu, Jiuxiang Rossi, Ryan A. Chen, Changyou Sun, Tong Computer Vision and Pattern Recognition Multimodal large language models (MLLMs) have recently shown great progress in text-rich image understanding, yet they still struggle with complex, multi-page visually-rich documents. Traditional methods using document parsers for retrieval-augmented generation suffer from performance and efficiency limitations, while directly presenting all pages to MLLMs leads to inefficiencies, especially with lengthy ones. In this work, we present a novel framework named **S**elf-**V**isual **R**etrieval-**A**ugmented **G**eneration (SV-RAG), which can broaden horizons of any MLLM to support long-document understanding. We demonstrate that **MLLMs themselves can be an effective multimodal retriever** to fetch relevant pages and then answer user questions based on these pages. SV-RAG is implemented with two specific MLLM adapters, one for evidence page retrieval and the other for question answering. Empirical results show state-of-the-art performance on public benchmarks, demonstrating the effectiveness of SV-RAG. |
| title | SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2411.01106 |