MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hsiao, Chi-Hsiang, Wang, Yi-Cheng, Lin, Tzung-Sheng, Yeh, Yi-Ren, Chen, Chu-Song
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918454285565952
author Hsiao, Chi-Hsiang
Wang, Yi-Cheng
Lin, Tzung-Sheng
Yeh, Yi-Ren
Chen, Chu-Song
author_facet Hsiao, Chi-Hsiang
Wang, Yi-Cheng
Lin, Tzung-Sheng
Yeh, Yi-Ren
Chen, Chu-Song
contents Retrieval-augmented generation (RAG) enables large language models (LLMs) to dynamically access external information, which is powerful for answering questions over previously unseen documents. Nonetheless, they struggle with high-level conceptual understanding and holistic comprehension due to limited context windows, which constrain their ability to perform deep reasoning over long-form, domain-specific content such as full-length books. To solve this problem, knowledge graphs (KGs) have been leveraged to provide entity-centric structure and hierarchical summaries, offering more structured support for reasoning. However, existing KG-based RAG solutions remain restricted to text-only inputs and fail to leverage the complementary insights provided by other modalities such as vision. On the other hand, reasoning from visual documents requires textual, visual, and spatial cues into structured, hierarchical concepts. To address this issue, we introduce a multimodal knowledge graph-based RAG that enables cross-modal reasoning for better content understanding. Our method incorporates visual cues into the construction of knowledge graphs, the retrieval phase, and the answer generation process. Experimental results across both global and fine-grained question answering tasks show that our approach consistently outperforms existing RAG-based approaches on both textual and multimodal corpora.
format Preprint
id arxiv_https___arxiv_org_abs_2512_20626
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
Hsiao, Chi-Hsiang
Wang, Yi-Cheng
Lin, Tzung-Sheng
Yeh, Yi-Ren
Chen, Chu-Song
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Information Retrieval
Retrieval-augmented generation (RAG) enables large language models (LLMs) to dynamically access external information, which is powerful for answering questions over previously unseen documents. Nonetheless, they struggle with high-level conceptual understanding and holistic comprehension due to limited context windows, which constrain their ability to perform deep reasoning over long-form, domain-specific content such as full-length books. To solve this problem, knowledge graphs (KGs) have been leveraged to provide entity-centric structure and hierarchical summaries, offering more structured support for reasoning. However, existing KG-based RAG solutions remain restricted to text-only inputs and fail to leverage the complementary insights provided by other modalities such as vision. On the other hand, reasoning from visual documents requires textual, visual, and spatial cues into structured, hierarchical concepts. To address this issue, we introduce a multimodal knowledge graph-based RAG that enables cross-modal reasoning for better content understanding. Our method incorporates visual cues into the construction of knowledge graphs, the retrieval phase, and the answer generation process. Experimental results across both global and fine-grained question answering tasks show that our approach consistently outperforms existing RAG-based approaches on both textual and multimodal corpora.
title MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Information Retrieval
url https://arxiv.org/abs/2512.20626