Multimodal Hypothetical Summary for Retrieval-based Multi-image Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Peize, Si, Qingyi, Fu, Peng, Lin, Zheng, Wang, Yan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Object Attribute Matters in Visual Question Answering
von: Li, Peize, et al.
Veröffentlicht: (2023)
von: Li, Peize, et al.
Veröffentlicht: (2023)
Towards Flexible Evaluation for Generative Visual Question Answering
von: Ji, Huishan, et al.
Veröffentlicht: (2024)
von: Ji, Huishan, et al.
Veröffentlicht: (2024)
Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation
von: Xu, Quanxing, et al.
Veröffentlicht: (2026)
von: Xu, Quanxing, et al.
Veröffentlicht: (2026)
SnapNTell: Enhancing Entity-Centric Visual Question Answering with Retrieval Augmented Multimodal LLM
von: Qiu, Jielin, et al.
Veröffentlicht: (2024)
von: Qiu, Jielin, et al.
Veröffentlicht: (2024)
Object Retrieval for Visual Question Answering with Outside Knowledge
von: Kan, Shichao, et al.
Veröffentlicht: (2024)
von: Kan, Shichao, et al.
Veröffentlicht: (2024)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
Multimodal Rationales for Explainable Visual Question Answering
von: Li, Kun, et al.
Veröffentlicht: (2024)
von: Li, Kun, et al.
Veröffentlicht: (2024)
AVIR: Adaptive Visual In-Document Retrieval for Efficient Multi-Page Document Question Answering
von: Li, Zongmin, et al.
Veröffentlicht: (2026)
von: Li, Zongmin, et al.
Veröffentlicht: (2026)
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
von: Xue, Junxiao, et al.
Veröffentlicht: (2024)
von: Xue, Junxiao, et al.
Veröffentlicht: (2024)
PartRAG: Retrieval-Augmented Part-Level 3D Generation and Editing
von: Li, Peize, et al.
Veröffentlicht: (2026)
von: Li, Peize, et al.
Veröffentlicht: (2026)
VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering
von: Lim, Qi Zhi, et al.
Veröffentlicht: (2025)
von: Lim, Qi Zhi, et al.
Veröffentlicht: (2025)
MOTOR: Multimodal Optimal Transport via Grounded Retrieval in Medical Visual Question Answering
von: Shaaban, Mai A., et al.
Veröffentlicht: (2025)
von: Shaaban, Mai A., et al.
Veröffentlicht: (2025)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
von: Di, Shangzhe, et al.
Veröffentlicht: (2025)
von: Di, Shangzhe, et al.
Veröffentlicht: (2025)
MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering
von: Peng, Jingwei, et al.
Veröffentlicht: (2025)
von: Peng, Jingwei, et al.
Veröffentlicht: (2025)
AstroRAG -- A Pagerank-Based Retrieval-Augmented Generation Pipeline for Question Answering in Astronomy
von: Wang, Zhifeng, et al.
Veröffentlicht: (2026)
von: Wang, Zhifeng, et al.
Veröffentlicht: (2026)
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems
von: Xu, Quanxing, et al.
Veröffentlicht: (2026)
von: Xu, Quanxing, et al.
Veröffentlicht: (2026)
Relevance-aware Multi-context Contrastive Decoding for Retrieval-augmented Visual Question Answering
von: Kim, Jongha, et al.
Veröffentlicht: (2026)
von: Kim, Jongha, et al.
Veröffentlicht: (2026)
Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
von: Wang, Zining, et al.
Veröffentlicht: (2025)
von: Wang, Zining, et al.
Veröffentlicht: (2025)
Towards Top-Down Reasoning: An Explainable Multi-Agent Approach for Visual Question Answering
von: Wang, Zeqing, et al.
Veröffentlicht: (2023)
von: Wang, Zeqing, et al.
Veröffentlicht: (2023)
Design as Desired: Utilizing Visual Question Answering for Multimodal Pre-training
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
M$^3$-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering
von: Ma, Jiatong, et al.
Veröffentlicht: (2026)
von: Ma, Jiatong, et al.
Veröffentlicht: (2026)
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering
von: Lim, Su Hyeon, et al.
Veröffentlicht: (2024)
von: Lim, Su Hyeon, et al.
Veröffentlicht: (2024)
AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
von: Zeng, Kang, et al.
Veröffentlicht: (2025)
von: Zeng, Kang, et al.
Veröffentlicht: (2025)
LLaVA-ReID: Selective Multi-image Questioner for Interactive Person Re-Identification
von: Lu, Yiding, et al.
Veröffentlicht: (2025)
von: Lu, Yiding, et al.
Veröffentlicht: (2025)
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
von: Lee, Dosung, et al.
Veröffentlicht: (2025)
von: Lee, Dosung, et al.
Veröffentlicht: (2025)
Exploring Multimodal LMMs for Online Episodic Memory Question Answering on the Edge
von: Lando, Giuseppe, et al.
Veröffentlicht: (2026)
von: Lando, Giuseppe, et al.
Veröffentlicht: (2026)
Glance and Focus: Memory Prompting for Multi-Event Video Question Answering
von: Bai, Ziyi, et al.
Veröffentlicht: (2024)
von: Bai, Ziyi, et al.
Veröffentlicht: (2024)
Multimodal Integration of Human-Like Attention in Visual Question Answering
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
Benchmarking Large Multimodal Models for Ophthalmic Visual Question Answering with OphthalWeChat
von: Xu, Pusheng, et al.
Veröffentlicht: (2025)
von: Xu, Pusheng, et al.
Veröffentlicht: (2025)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
von: Zhang, Chengyi, et al.
Veröffentlicht: (2026)
von: Zhang, Chengyi, et al.
Veröffentlicht: (2026)
Improving Data Augmentation for Robust Visual Question Answering with Effective Curriculum Learning
von: Zheng, Yuhang, et al.
Veröffentlicht: (2024)
von: Zheng, Yuhang, et al.
Veröffentlicht: (2024)
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
von: Luo, Jingzhou, et al.
Veröffentlicht: (2025)
von: Luo, Jingzhou, et al.
Veröffentlicht: (2025)
Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling
von: Movva, Prahitha, et al.
Veröffentlicht: (2025)
von: Movva, Prahitha, et al.
Veröffentlicht: (2025)
RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering
von: Wang, Yuduo, et al.
Veröffentlicht: (2023)
von: Wang, Yuduo, et al.
Veröffentlicht: (2023)
AIM: Asymmetric Information Masking for Visual Question Answering Continual Learning
von: Zhang, Peifeng, et al.
Veröffentlicht: (2026)
von: Zhang, Peifeng, et al.
Veröffentlicht: (2026)
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
von: Song, Enxin, et al.
Veröffentlicht: (2024)
von: Song, Enxin, et al.
Veröffentlicht: (2024)
VDMA: Video Question Answering with Dynamically Generated Multi-Agents
von: Kugo, Noriyuki, et al.
Veröffentlicht: (2024)
von: Kugo, Noriyuki, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Object Attribute Matters in Visual Question Answering
von: Li, Peize, et al.
Veröffentlicht: (2023) -
Towards Flexible Evaluation for Generative Visual Question Answering
von: Ji, Huishan, et al.
Veröffentlicht: (2024) -
Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation
von: Xu, Quanxing, et al.
Veröffentlicht: (2026) -
SnapNTell: Enhancing Entity-Centric Visual Question Answering with Retrieval Augmented Multimodal LLM
von: Qiu, Jielin, et al.
Veröffentlicht: (2024) -
Object Retrieval for Visual Question Answering with Outside Knowledge
von: Kan, Shichao, et al.
Veröffentlicht: (2024)