Retrieval Meets Reasoning: Even High-school Textbook Knowledge Benefits Multimodal Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tan, Cheng, Wei, Jingxuan, Sun, Linzhuang, Gao, Zhangyang, Li, Siyuan, Yu, Bihui, Guo, Ruifeng, Li, Stan Z.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911896657985536
author Tan, Cheng
Wei, Jingxuan
Sun, Linzhuang
Gao, Zhangyang
Li, Siyuan
Yu, Bihui
Guo, Ruifeng
Li, Stan Z.
author_facet Tan, Cheng
Wei, Jingxuan
Sun, Linzhuang
Gao, Zhangyang
Li, Siyuan
Yu, Bihui
Guo, Ruifeng
Li, Stan Z.
contents Large language models equipped with retrieval-augmented generation (RAG) represent a burgeoning field aimed at enhancing answering capabilities by leveraging external knowledge bases. Although the application of RAG with language-only models has been extensively explored, its adaptation into multimodal vision-language models remains nascent. Going beyond mere answer generation, the primary goal of multimodal RAG is to cultivate the models' ability to reason in response to relevant queries. To this end, we introduce a novel multimodal RAG framework named RMR (Retrieval Meets Reasoning). The RMR framework employs a bi-modal retrieval module to identify the most relevant question-answer pairs, which then serve as scaffolds for the multimodal reasoning process. This training-free approach not only encourages the model to engage deeply with the reasoning processes inherent in the retrieved content but also facilitates the generation of answers that are precise and richly interpretable. Surprisingly, utilizing solely the ScienceQA dataset, collected from elementary and high school science curricula, RMR significantly boosts the performance of various vision-language models across a spectrum of benchmark datasets, including A-OKVQA, MMBench, and SEED. These outcomes highlight the substantial potential of our multimodal retrieval and reasoning mechanism to improve the reasoning capabilities of vision-language models.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20834
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Retrieval Meets Reasoning: Even High-school Textbook Knowledge Benefits Multimodal Reasoning
Tan, Cheng
Wei, Jingxuan
Sun, Linzhuang
Gao, Zhangyang
Li, Siyuan
Yu, Bihui
Guo, Ruifeng
Li, Stan Z.
Computer Vision and Pattern Recognition
Large language models equipped with retrieval-augmented generation (RAG) represent a burgeoning field aimed at enhancing answering capabilities by leveraging external knowledge bases. Although the application of RAG with language-only models has been extensively explored, its adaptation into multimodal vision-language models remains nascent. Going beyond mere answer generation, the primary goal of multimodal RAG is to cultivate the models' ability to reason in response to relevant queries. To this end, we introduce a novel multimodal RAG framework named RMR (Retrieval Meets Reasoning). The RMR framework employs a bi-modal retrieval module to identify the most relevant question-answer pairs, which then serve as scaffolds for the multimodal reasoning process. This training-free approach not only encourages the model to engage deeply with the reasoning processes inherent in the retrieved content but also facilitates the generation of answers that are precise and richly interpretable. Surprisingly, utilizing solely the ScienceQA dataset, collected from elementary and high school science curricula, RMR significantly boosts the performance of various vision-language models across a spectrum of benchmark datasets, including A-OKVQA, MMBench, and SEED. These outcomes highlight the substantial potential of our multimodal retrieval and reasoning mechanism to improve the reasoning capabilities of vision-language models.
title Retrieval Meets Reasoning: Even High-school Textbook Knowledge Benefits Multimodal Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.20834