MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Gong, Ziyu, Mai, Chengcheng, Huang, Yihua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset
by: Tran, Quang-Linh, et al.
Published: (2025)
by: Tran, Quang-Linh, et al.
Published: (2025)
Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
by: Zhou, Ao, et al.
Published: (2025)
by: Zhou, Ao, et al.
Published: (2025)
RAG-VisualRec: An Open Resource for Vision- and Text-Enhanced Retrieval-Augmented Generation in Recommendation
by: Tourani, Ali, et al.
Published: (2025)
by: Tourani, Ali, et al.
Published: (2025)
Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering
by: Koh, Junyoung, et al.
Published: (2026)
by: Koh, Junyoung, et al.
Published: (2026)
AsCL: An Asymmetry-sensitive Contrastive Learning Method for Image-Text Retrieval with Cross-Modal Fusion
by: Gong, Ziyu, et al.
Published: (2024)
by: Gong, Ziyu, et al.
Published: (2024)
Agentic Mixed-Source Multi-Modal Misinformation Detection with Adaptive Test-Time Scaling
by: Jiang, Wei, et al.
Published: (2026)
by: Jiang, Wei, et al.
Published: (2026)
StePO-Rec: Towards Personalized Outfit Styling Assistant via Knowledge-Guided Multi-Step Reasoning
by: Bi, Yuxi, et al.
Published: (2025)
by: Bi, Yuxi, et al.
Published: (2025)
A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task
by: Deng, Jiaqi, et al.
Published: (2025)
by: Deng, Jiaqi, et al.
Published: (2025)
Leveraging Weak Cross-Modal Guidance for Coherence Modelling via Iterative Learning
by: Bin, Yi, et al.
Published: (2024)
by: Bin, Yi, et al.
Published: (2024)
MedCoT-RAG: Causal Chain-of-Thought RAG for Medical Question Answering
by: Wang, Ziyu, et al.
Published: (2025)
by: Wang, Ziyu, et al.
Published: (2025)
Modality-Aware Identity Construction and Counterfactual Structure Learning for ID-Free Multimodal Recommendation
by: Ma, Hongjian, et al.
Published: (2026)
by: Ma, Hongjian, et al.
Published: (2026)
Adaptive Multi-Agent Reasoning for Text-to-Video Retrieval
by: Wu, Jiaxin, et al.
Published: (2025)
by: Wu, Jiaxin, et al.
Published: (2025)
Mitigating Modality Bias in Multi-modal Entity Alignment from a Causal Perspective
by: Su, Taoyu, et al.
Published: (2025)
by: Su, Taoyu, et al.
Published: (2025)
Ambiguity-Aware and High-Order Relation Learning for Multi-Grained Image-Text Matching
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions
by: Wang, Tianshi, et al.
Published: (2023)
by: Wang, Tianshi, et al.
Published: (2023)
MDF: A Dynamic Fusion Model for Multi-modal Fake News Detection
by: Lv, Hongzhen, et al.
Published: (2024)
by: Lv, Hongzhen, et al.
Published: (2024)
KET-RAG: A Cost-Efficient Multi-Granular Indexing Framework for Graph-RAG
by: Huang, Yiqian, et al.
Published: (2025)
by: Huang, Yiqian, et al.
Published: (2025)
CLEAR: Null-Space Projection for Cross-Modal De-Redundancy in Multimodal Recommendation
by: Zhan, Hao, et al.
Published: (2026)
by: Zhan, Hao, et al.
Published: (2026)
Spectrum-based Modality Representation Fusion Graph Convolutional Network for Multimodal Recommendation
by: Ong, Rongqing Kenneth, et al.
Published: (2024)
by: Ong, Rongqing Kenneth, et al.
Published: (2024)
U-Sticker: A Large-Scale Multi-Domain User Sticker Dataset for Retrieval and Personalization
by: Chee, Heng Er Metilda, et al.
Published: (2025)
by: Chee, Heng Er Metilda, et al.
Published: (2025)
Multimodal Graph Neural Network for Recommendation with Dynamic De-redundancy and Modality-Guided Feature De-noisy
by: Mo, Feng, et al.
Published: (2024)
by: Mo, Feng, et al.
Published: (2024)
Socially Aware Music Recommendation: A Multi-Modal Graph Neural Networks for Collaborative Music Consumption and Community-Based Engagement
by: Ziaoddini, Kajwan
Published: (2025)
by: Ziaoddini, Kajwan
Published: (2025)
Enabling Collaborative Parametric Knowledge Calibration for Retrieval-Augmented Vision Question Answering
by: Deng, Jiaqi, et al.
Published: (2025)
by: Deng, Jiaqi, et al.
Published: (2025)
SoccerRAG: Multimodal Soccer Information Retrieval via Natural Queries
by: Strand, Aleksander Theo, et al.
Published: (2024)
by: Strand, Aleksander Theo, et al.
Published: (2024)
Demo: Soccer Information Retrieval via Natural Queries using SoccerRAG
by: Strand, Aleksander Theo, et al.
Published: (2024)
by: Strand, Aleksander Theo, et al.
Published: (2024)
Improving the Consistency in Cross-Lingual Cross-Modal Retrieval with 1-to-K Contrastive Learning
by: Nie, Zhijie, et al.
Published: (2024)
by: Nie, Zhijie, et al.
Published: (2024)
Multimodal Pre-training Framework for Sequential Recommendation via Contrastive Learning
by: Zhang, Lingzi, et al.
Published: (2023)
by: Zhang, Lingzi, et al.
Published: (2023)
Why Multi-Interest Fairness Matters: Hypergraph Contrastive Multi-Interest Learning for Fair Conversational Recommender System
by: Zheng, Yongsen, et al.
Published: (2025)
by: Zheng, Yongsen, et al.
Published: (2025)
StepChain GraphRAG: Reasoning Over Knowledge Graphs for Multi-Hop Question Answering
by: Ni, Tengjun, et al.
Published: (2025)
by: Ni, Tengjun, et al.
Published: (2025)
Don't Lose Yourself: Boosting Multimodal Recommendation via Reducing Node-neighbor Discrepancy in Graph Convolutional Network
by: Chen, Zheyu, et al.
Published: (2024)
by: Chen, Zheyu, et al.
Published: (2024)
A Multi-Source Retrieval Question Answering Framework Based on RAG
by: Wu, Ridong, et al.
Published: (2024)
by: Wu, Ridong, et al.
Published: (2024)
Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and Beyond
by: Wei, Tianxin, et al.
Published: (2024)
by: Wei, Tianxin, et al.
Published: (2024)
Uni-Retrieval: A Multi-Style Retrieval Framework for STEM's Education
by: Jia, Yanhao, et al.
Published: (2025)
by: Jia, Yanhao, et al.
Published: (2025)
ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective Reasoning
by: Luo, Pengfei, et al.
Published: (2025)
by: Luo, Pengfei, et al.
Published: (2025)
VeritasFi: An Adaptable, Multi-tiered RAG Framework for Multi-modal Financial Question Answering
by: Tai, Zhenghan, et al.
Published: (2025)
by: Tai, Zhenghan, et al.
Published: (2025)
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs
by: Li, Haoxuan, et al.
Published: (2025)
by: Li, Haoxuan, et al.
Published: (2025)
Unraveling Movie Genres through Cross-Attention Fusion of Bi-Modal Synergy of Poster
by: Nareti, Utsav Kumar, et al.
Published: (2024)
by: Nareti, Utsav Kumar, et al.
Published: (2024)
QuMAB: Query-based Multi-Annotator Behavior Modeling with Reliability under Sparse Labels
by: Zhang, Liyun, et al.
Published: (2025)
by: Zhang, Liyun, et al.
Published: (2025)
Decomposing Retrieval Failures in RAG for Long-Document Financial Question Answering
by: Kobeissi, Amine, et al.
Published: (2026)
by: Kobeissi, Amine, et al.
Published: (2026)
Learning Item Representations Directly from Multimodal Features for Effective Recommendation
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
Similar Items
-
OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset
by: Tran, Quang-Linh, et al.
Published: (2025) -
Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
by: Zhou, Ao, et al.
Published: (2025) -
RAG-VisualRec: An Open Resource for Vision- and Text-Enhanced Retrieval-Augmented Generation in Recommendation
by: Tourani, Ali, et al.
Published: (2025) -
Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering
by: Koh, Junyoung, et al.
Published: (2026) -
AsCL: An Asymmetry-sensitive Contrastive Learning Method for Image-Text Retrieval with Cross-Modal Fusion
by: Gong, Ziyu, et al.
Published: (2024)