Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xue, Junxiao, Deng, Quan, Yu, Fei, Wang, Yanhao, Wang, Jun, Li, Yuehua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning
von: Xue, Junxiao, et al.
Veröffentlicht: (2026)
von: Xue, Junxiao, et al.
Veröffentlicht: (2026)
InfoSyncNet: Information Synchronization Temporal Convolutional Network for Visual Speech Recognition
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
SnapNTell: Enhancing Entity-Centric Visual Question Answering with Retrieval Augmented Multimodal LLM
von: Qiu, Jielin, et al.
Veröffentlicht: (2024)
von: Qiu, Jielin, et al.
Veröffentlicht: (2024)
Multimodal Rationales for Explainable Visual Question Answering
von: Li, Kun, et al.
Veröffentlicht: (2024)
von: Li, Kun, et al.
Veröffentlicht: (2024)
Design as Desired: Utilizing Visual Question Answering for Multimodal Pre-training
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning
von: Yu, Fei, et al.
Veröffentlicht: (2025)
von: Yu, Fei, et al.
Veröffentlicht: (2025)
Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling
von: Movva, Prahitha, et al.
Veröffentlicht: (2025)
von: Movva, Prahitha, et al.
Veröffentlicht: (2025)
Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation
von: Xu, Quanxing, et al.
Veröffentlicht: (2026)
von: Xu, Quanxing, et al.
Veröffentlicht: (2026)
Enhancing Visual Question Answering through Ranking-Based Hybrid Training and Multimodal Fusion
von: Chen, Peiyuan, et al.
Veröffentlicht: (2024)
von: Chen, Peiyuan, et al.
Veröffentlicht: (2024)
A Visual Question Answering Method for SAR Ship: Breaking the Requirement for Multimodal Dataset Construction and Model Fine-Tuning
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
von: Lee, Dosung, et al.
Veröffentlicht: (2025)
von: Lee, Dosung, et al.
Veröffentlicht: (2025)
RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering
von: Wang, Yuduo, et al.
Veröffentlicht: (2023)
von: Wang, Yuduo, et al.
Veröffentlicht: (2023)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
von: Zhang, Chengyi, et al.
Veröffentlicht: (2026)
von: Zhang, Chengyi, et al.
Veröffentlicht: (2026)
VQ-VA World: Towards High-Quality Visual Question-Visual Answering
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
Visually Interpretable Subtask Reasoning for Visual Question Answering
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
A Comprehensive Evaluation of Transformer-Based Question Answering Models and RAG-Enhanced Design
von: Zhang, Zichen, et al.
Veröffentlicht: (2025)
von: Zhang, Zichen, et al.
Veröffentlicht: (2025)
AstroRAG -- A Pagerank-Based Retrieval-Augmented Generation Pipeline for Question Answering in Astronomy
von: Wang, Zhifeng, et al.
Veröffentlicht: (2026)
von: Wang, Zhifeng, et al.
Veröffentlicht: (2026)
Improving Few-Shot Change Detection Visual Question Answering via Decision-Ambiguity-guided Reinforcement Fine-Tuning
von: Dong, Fuyu, et al.
Veröffentlicht: (2025)
von: Dong, Fuyu, et al.
Veröffentlicht: (2025)
MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering
von: Peng, Jingwei, et al.
Veröffentlicht: (2025)
von: Peng, Jingwei, et al.
Veröffentlicht: (2025)
Multimodal Integration of Human-Like Attention in Visual Question Answering
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering
von: Chen, Zhuohong, et al.
Veröffentlicht: (2026)
von: Chen, Zhuohong, et al.
Veröffentlicht: (2026)
Object Retrieval for Visual Question Answering with Outside Knowledge
von: Kan, Shichao, et al.
Veröffentlicht: (2024)
von: Kan, Shichao, et al.
Veröffentlicht: (2024)
Multimodal Hypothetical Summary for Retrieval-based Multi-image Question Answering
von: Li, Peize, et al.
Veröffentlicht: (2024)
von: Li, Peize, et al.
Veröffentlicht: (2024)
Concept-Enhanced Multimodal RAG: Towards Interpretable and Accurate Radiology Report Generation
von: Salmè, Marco, et al.
Veröffentlicht: (2026)
von: Salmè, Marco, et al.
Veröffentlicht: (2026)
Reconstruction as a Bridge for Event-Based Visual Question Answering
von: Lou, Hanyue, et al.
Veröffentlicht: (2025)
von: Lou, Hanyue, et al.
Veröffentlicht: (2025)
Structure Causal Models and LLMs Integration in Medical Visual Question Answering
von: Xu, Zibo, et al.
Veröffentlicht: (2025)
von: Xu, Zibo, et al.
Veröffentlicht: (2025)
Benchmarking Multimodal RAG through a Chart-based Document Question-Answering Generation Framework
von: Yang, Yuming, et al.
Veröffentlicht: (2025)
von: Yang, Yuming, et al.
Veröffentlicht: (2025)
Object Attribute Matters in Visual Question Answering
von: Li, Peize, et al.
Veröffentlicht: (2023)
von: Li, Peize, et al.
Veröffentlicht: (2023)
VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
von: Li, Yuyi, et al.
Veröffentlicht: (2025)
von: Li, Yuyi, et al.
Veröffentlicht: (2025)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
Selectively Answering Visual Questions
von: Eisenschlos, Julian Martin, et al.
Veröffentlicht: (2024)
von: Eisenschlos, Julian Martin, et al.
Veröffentlicht: (2024)
A Simple LLM Framework for Long-Range Video Question-Answering
von: Zhang, Ce, et al.
Veröffentlicht: (2023)
von: Zhang, Ce, et al.
Veröffentlicht: (2023)
Towards Comprehensive Interactive Change Understanding in Remote Sensing: A Large-scale Dataset and Dual-granularity Enhanced VLM
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
A Trustworthy Method for Multimodal Emotion Recognition
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
Visual Question Answering Instruction: Unlocking Multimodal Large Language Model To Domain-Specific Visual Multitasks
von: Lee, Jusung, et al.
Veröffentlicht: (2024)
von: Lee, Jusung, et al.
Veröffentlicht: (2024)
MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
von: Song, Seokwon, et al.
Veröffentlicht: (2025)
von: Song, Seokwon, et al.
Veröffentlicht: (2025)
RoadscapesQA: A Multitask, Multimodal Dataset for Visual Question Answering on Indian Roads
von: Iyer, Vijayasri, et al.
Veröffentlicht: (2026)
von: Iyer, Vijayasri, et al.
Veröffentlicht: (2026)
Targeted Visual Prompting for Medical Visual Question Answering
von: Tascon-Morales, Sergio, et al.
Veröffentlicht: (2024)
von: Tascon-Morales, Sergio, et al.
Veröffentlicht: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning
von: Xue, Junxiao, et al.
Veröffentlicht: (2026) -
InfoSyncNet: Information Synchronization Temporal Convolutional Network for Visual Speech Recognition
von: Xue, Junxiao, et al.
Veröffentlicht: (2025) -
SnapNTell: Enhancing Entity-Centric Visual Question Answering with Retrieval Augmented Multimodal LLM
von: Qiu, Jielin, et al.
Veröffentlicht: (2024) -
Multimodal Rationales for Explainable Visual Question Answering
von: Li, Kun, et al.
Veröffentlicht: (2024) -
Design as Desired: Utilizing Visual Question Answering for Multimodal Pre-training
von: Su, Tongkun, et al.
Veröffentlicht: (2024)