VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Suri, Manan, Mathur, Puneet, Dernoncourt, Franck, Goswami, Kanika, Rossi, Ryan A., Manocha, Dinesh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PlotGen: Multi-Agent LLM-based Scientific Data Visualization via Multimodal Feedback
by: Goswami, Kanika, et al.
Published: (2025)
by: Goswami, Kanika, et al.
Published: (2025)
ChartLens: Fine-grained Visual Attribution in Charts
by: Suri, Manan, et al.
Published: (2025)
by: Suri, Manan, et al.
Published: (2025)
ChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution
by: Goswami, Kanika, et al.
Published: (2025)
by: Goswami, Kanika, et al.
Published: (2025)
PlotEdit: Natural Language-Driven Accessible Chart Editing in PDFs via Multimodal LLM Agents
by: Goswami, Kanika, et al.
Published: (2025)
by: Goswami, Kanika, et al.
Published: (2025)
Structured Uncertainty guided Clarification for LLM Agents
by: Suri, Manan, et al.
Published: (2025)
by: Suri, Manan, et al.
Published: (2025)
Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents
by: Suri, Manan, et al.
Published: (2025)
by: Suri, Manan, et al.
Published: (2025)
DocEdit-v2: Document Structure Editing Via Multimodal LLM Grounding
by: Suri, Manan, et al.
Published: (2024)
by: Suri, Manan, et al.
Published: (2024)
DIAGRAMS: A Review Framework for Reasoning-Level Attribution in Diagram QA
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2026)
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2026)
DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2026)
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2026)
Learning Illumination Control in Diffusion Models
by: Anand, Nishit, et al.
Published: (2026)
by: Anand, Nishit, et al.
Published: (2026)
Charts Are Not Images: On the Challenges of Scientific Chart Editing
by: Li, Shawn, et al.
Published: (2025)
by: Li, Shawn, et al.
Published: (2025)
Cluster-R1: Large Reasoning Models Are Instruction-following Clustering Agents
by: Qing, Peijun, et al.
Published: (2026)
by: Qing, Peijun, et al.
Published: (2026)
FASH-iCNN: Making Editorial Fashion Identity Inspectable Through Multimodal CNN Probing
by: Adeyemi, Morayo Danielle, et al.
Published: (2026)
by: Adeyemi, Morayo Danielle, et al.
Published: (2026)
MODS: Moderating a Mixture of Document Speakers to Summarize Debatable Queries in Document Collections
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
by: Chen, Jian, et al.
Published: (2025)
by: Chen, Jian, et al.
Published: (2025)
Document Attribution: Examining Citation Relationships using Large Language Models
by: Rawte, Vipula, et al.
Published: (2025)
by: Rawte, Vipula, et al.
Published: (2025)
Blind to the Human Touch: Overlap Bias in LLM-Based Summary Evaluation
by: Fang, Jiangnan, et al.
Published: (2026)
by: Fang, Jiangnan, et al.
Published: (2026)
Test-Time Strategies for More Efficient and Accurate Agentic RAG
by: Zhang, Brian, et al.
Published: (2026)
by: Zhang, Brian, et al.
Published: (2026)
Trust but Verify: Introducing DAVinCI -- A Framework for Dual Attribution and Verification in Claim Inference for Language Models
by: Rawte, Vipula, et al.
Published: (2026)
by: Rawte, Vipula, et al.
Published: (2026)
Retrieval Improvements Do Not Guarantee Better Answers: A Study of RAG for AI Policy QA
by: Mathur, Saahil, et al.
Published: (2026)
by: Mathur, Saahil, et al.
Published: (2026)
Local Hybrid Retrieval-Augmented Document QA
by: Astrino, Paolo
Published: (2025)
by: Astrino, Paolo
Published: (2025)
Multi-LLM QA with Embodied Exploration
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
RECAP: Retrieval-Augmented Audio Captioning
by: Ghosh, Sreyan, et al.
Published: (2023)
by: Ghosh, Sreyan, et al.
Published: (2023)
Benchmarking Knowledge-Extraction Attack and Defense on Retrieval-Augmented Generation
by: Qi, Zhisheng, et al.
Published: (2026)
by: Qi, Zhisheng, et al.
Published: (2026)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
by: Tanaka, Ryota, et al.
Published: (2025)
by: Tanaka, Ryota, et al.
Published: (2025)
Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering
by: Ngo, Nghia Trung, et al.
Published: (2024)
by: Ngo, Nghia Trung, et al.
Published: (2024)
GRS-QA -- Graph Reasoning-Structured Question Answering Dataset
by: Pahilajani, Anish, et al.
Published: (2024)
by: Pahilajani, Anish, et al.
Published: (2024)
EM-GANSim: Real-time and Accurate EM Simulation Using Conditional GANs for 3D Indoor Scenes
by: Wang, Ruichen, et al.
Published: (2024)
by: Wang, Ruichen, et al.
Published: (2024)
HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents
by: Tong, Anyang, et al.
Published: (2025)
by: Tong, Anyang, et al.
Published: (2025)
Do Audio-Visual Large Language Models Really See and Hear?
by: Selvakumar, Ramaneswaran, et al.
Published: (2026)
by: Selvakumar, Ramaneswaran, et al.
Published: (2026)
MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces
by: Luera, Reuben A., et al.
Published: (2025)
by: Luera, Reuben A., et al.
Published: (2025)
DocuBits: VR Document Decomposition for Procedural Task Completion
by: Lee, Geonsun, et al.
Published: (2024)
by: Lee, Geonsun, et al.
Published: (2024)
MemCtrl: Using MLLMs as Active Memory Controllers on Embodied Agents
by: Dorbala, Vishnu Sashank, et al.
Published: (2026)
by: Dorbala, Vishnu Sashank, et al.
Published: (2026)
DynaSaur: Large Language Agents Beyond Predefined Actions
by: Nguyen, Dang, et al.
Published: (2024)
by: Nguyen, Dang, et al.
Published: (2024)
BoMuDANet: Unsupervised Adaptation for Visual Scene Understanding in Unstructured Driving Environments
by: Kothandaraman, Divya, et al.
Published: (2020)
by: Kothandaraman, Divya, et al.
Published: (2020)
Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA
by: Pu, Yuan, et al.
Published: (2024)
by: Pu, Yuan, et al.
Published: (2024)
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
by: Zhao, Baoquan, et al.
Published: (2025)
by: Zhao, Baoquan, et al.
Published: (2025)
Retrieval Augmented Generation for Domain-specific Question Answering
by: Sharma, Sanat, et al.
Published: (2024)
by: Sharma, Sanat, et al.
Published: (2024)
Mixture of Structural-and-Textual Retrieval over Text-rich Graph Knowledge Bases
by: Lei, Yongjia, et al.
Published: (2025)
by: Lei, Yongjia, et al.
Published: (2025)
Anticipatory Planning for Multimodal AI Agents
by: Liang, Yongyuan, et al.
Published: (2026)
by: Liang, Yongyuan, et al.
Published: (2026)
Similar Items
-
PlotGen: Multi-Agent LLM-based Scientific Data Visualization via Multimodal Feedback
by: Goswami, Kanika, et al.
Published: (2025) -
ChartLens: Fine-grained Visual Attribution in Charts
by: Suri, Manan, et al.
Published: (2025) -
ChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution
by: Goswami, Kanika, et al.
Published: (2025) -
PlotEdit: Natural Language-Driven Accessible Chart Editing in PDFs via Multimodal LLM Agents
by: Goswami, Kanika, et al.
Published: (2025) -
Structured Uncertainty guided Clarification for LLM Agents
by: Suri, Manan, et al.
Published: (2025)