A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Jiaqi, Wu, Zonghan, Huo, Huan, Xu, Guandong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enabling Collaborative Parametric Knowledge Calibration for Retrieval-Augmented Vision Question Answering
by: Deng, Jiaqi, et al.
Published: (2025)
by: Deng, Jiaqi, et al.
Published: (2025)
From Swath to Full-Disc: Advancing Precipitation Retrieval with Multimodal Knowledge Expansion
by: Wang, Zheng, et al.
Published: (2025)
by: Wang, Zheng, et al.
Published: (2025)
A Comprehensive Survey on Composed Image Retrieval
by: Song, Xuemeng, et al.
Published: (2025)
by: Song, Xuemeng, et al.
Published: (2025)
MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning
by: Gong, Ziyu, et al.
Published: (2025)
by: Gong, Ziyu, et al.
Published: (2025)
Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
by: Zhou, Ao, et al.
Published: (2025)
by: Zhou, Ao, et al.
Published: (2025)
A Unified Optimal Transport Framework for Cross-Modal Retrieval with Noisy Labels
by: Han, Haochen, et al.
Published: (2024)
by: Han, Haochen, et al.
Published: (2024)
UCDR-Adapter: Exploring Adaptation of Pre-Trained Vision-Language Models for Universal Cross-Domain Retrieval
by: Jiang, Haoyu, et al.
Published: (2024)
by: Jiang, Haoyu, et al.
Published: (2024)
VisTopics: A Visual Semantic Unsupervised Approach to Topic Modeling of Video and Image Data
by: Lokmanoglu, Ayse D, et al.
Published: (2025)
by: Lokmanoglu, Ayse D, et al.
Published: (2025)
Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning
by: Luo, Tianci, et al.
Published: (2026)
by: Luo, Tianci, et al.
Published: (2026)
OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset
by: Tran, Quang-Linh, et al.
Published: (2025)
by: Tran, Quang-Linh, et al.
Published: (2025)
Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues
by: Fu, Junchen, et al.
Published: (2026)
by: Fu, Junchen, et al.
Published: (2026)
CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection
by: Li, Fanxiao, et al.
Published: (2025)
by: Li, Fanxiao, et al.
Published: (2025)
NativE: Multi-modal Knowledge Graph Completion in the Wild
by: Zhang, Yichi, et al.
Published: (2024)
by: Zhang, Yichi, et al.
Published: (2024)
Self-distilled Dynamic Fusion Network for Language-based Fashion Retrieval
by: Wu, Yiming, et al.
Published: (2024)
by: Wu, Yiming, et al.
Published: (2024)
Breaking the Curse of Knowledge: Towards Effective Multimodal Recommendation using Knowledge Soft Integration
by: Ouyang, Kai, et al.
Published: (2023)
by: Ouyang, Kai, et al.
Published: (2023)
Interactive Multi-Turn Retrieval for Health Videos
by: Wu, Chengzheng, et al.
Published: (2026)
by: Wu, Chengzheng, et al.
Published: (2026)
StePO-Rec: Towards Personalized Outfit Styling Assistant via Knowledge-Guided Multi-Step Reasoning
by: Bi, Yuxi, et al.
Published: (2025)
by: Bi, Yuxi, et al.
Published: (2025)
DSRAG: A Domain-Specific Retrieval Framework Based on Document-derived Multimodal Knowledge Graph
by: Yang, Mengzheng, et al.
Published: (2025)
by: Yang, Mengzheng, et al.
Published: (2025)
WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering
by: Zhu, Yingjian, et al.
Published: (2026)
by: Zhu, Yingjian, et al.
Published: (2026)
PhotoBench: Beyond Visual Matching Towards Personalized Intent-Driven Photo Retrieval
by: Xu, Tianyi, et al.
Published: (2026)
by: Xu, Tianyi, et al.
Published: (2026)
Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
by: He, Xiao, et al.
Published: (2025)
by: He, Xiao, et al.
Published: (2025)
Knowledge-aware Diffusion-Enhanced Multimedia Recommendation
by: Mo, Xian, et al.
Published: (2025)
by: Mo, Xian, et al.
Published: (2025)
A Survey of Multimodal Composite Editing and Retrieval
by: Li, Suyan, et al.
Published: (2024)
by: Li, Suyan, et al.
Published: (2024)
Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering
by: Koh, Junyoung, et al.
Published: (2026)
by: Koh, Junyoung, et al.
Published: (2026)
Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis
by: Pegia, Maria-Eirini, et al.
Published: (2026)
by: Pegia, Maria-Eirini, et al.
Published: (2026)
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
by: Ning, Hailong, et al.
Published: (2025)
by: Ning, Hailong, et al.
Published: (2025)
HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
by: Li, Jun, et al.
Published: (2025)
by: Li, Jun, et al.
Published: (2025)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
by: Xiao, Jian, et al.
Published: (2025)
by: Xiao, Jian, et al.
Published: (2025)
Joint-Dataset Learning and Cross-Consistent Regularization for Text-to-Motion Retrieval
by: Messina, Nicola, et al.
Published: (2024)
by: Messina, Nicola, et al.
Published: (2024)
Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-Identification
by: Zhang, Pingping, et al.
Published: (2024)
by: Zhang, Pingping, et al.
Published: (2024)
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
by: Liu, Han, et al.
Published: (2025)
by: Liu, Han, et al.
Published: (2025)
Towards Identity-Aware Cross-Modal Retrieval: a Dataset and a Baseline
by: Messina, Nicola, et al.
Published: (2024)
by: Messina, Nicola, et al.
Published: (2024)
Efficient Self-Supervised Video Hashing with Selective State Spaces
by: Wang, Jinpeng, et al.
Published: (2024)
by: Wang, Jinpeng, et al.
Published: (2024)
Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval
by: Li, Jun, et al.
Published: (2026)
by: Li, Jun, et al.
Published: (2026)
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
by: Kong, Fanheng, et al.
Published: (2025)
by: Kong, Fanheng, et al.
Published: (2025)
Learning Partially-Decorrelated Common Spaces for Ad-hoc Video Search
by: Hu, Fan, et al.
Published: (2025)
by: Hu, Fan, et al.
Published: (2025)
Ambiguity-Aware and High-Order Relation Learning for Multi-Grained Image-Text Matching
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos
by: Li, Minghan, et al.
Published: (2026)
by: Li, Minghan, et al.
Published: (2026)
AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing
by: Lian, Niu, et al.
Published: (2025)
by: Lian, Niu, et al.
Published: (2025)
Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval
by: Xiao, Jian, et al.
Published: (2024)
by: Xiao, Jian, et al.
Published: (2024)
Similar Items
-
Enabling Collaborative Parametric Knowledge Calibration for Retrieval-Augmented Vision Question Answering
by: Deng, Jiaqi, et al.
Published: (2025) -
From Swath to Full-Disc: Advancing Precipitation Retrieval with Multimodal Knowledge Expansion
by: Wang, Zheng, et al.
Published: (2025) -
A Comprehensive Survey on Composed Image Retrieval
by: Song, Xuemeng, et al.
Published: (2025) -
MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning
by: Gong, Ziyu, et al.
Published: (2025) -
Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
by: Zhou, Ao, et al.
Published: (2025)