Combining Knowledge Graph and LLMs for Enhanced Zero-shot Visual Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tao, Qian, Fan, Xiaoyang, Xu, Yong, Zhu, Xingquan, Tang, Yufei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Perceiving Small Visual Details in Zero-shot Visual Question Answering with Multimodal LLMs
von: Zhang, Jiarui, et al.
Veröffentlicht: (2023)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2023)
Knowledge-Enhanced Dual-stream Zero-shot Composed Image Retrieval
von: Suo, Yucheng, et al.
Veröffentlicht: (2024)
von: Suo, Yucheng, et al.
Veröffentlicht: (2024)
Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation
von: Xu, Quanxing, et al.
Veröffentlicht: (2026)
von: Xu, Quanxing, et al.
Veröffentlicht: (2026)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
von: Romero, David, et al.
Veröffentlicht: (2024)
von: Romero, David, et al.
Veröffentlicht: (2024)
Map-based Modular Approach for Zero-shot Embodied Question Answering
von: Sakamoto, Koya, et al.
Veröffentlicht: (2024)
von: Sakamoto, Koya, et al.
Veröffentlicht: (2024)
Structure Causal Models and LLMs Integration in Medical Visual Question Answering
von: Xu, Zibo, et al.
Veröffentlicht: (2025)
von: Xu, Zibo, et al.
Veröffentlicht: (2025)
Object Retrieval for Visual Question Answering with Outside Knowledge
von: Kan, Shichao, et al.
Veröffentlicht: (2024)
von: Kan, Shichao, et al.
Veröffentlicht: (2024)
Zero-Shot Anomaly Detection in Battery Thermal Images Using Visual Question Answering with Prior Knowledge
von: Astrid, Marcella, et al.
Veröffentlicht: (2025)
von: Astrid, Marcella, et al.
Veröffentlicht: (2025)
KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering
von: Li, Zhiyang, et al.
Veröffentlicht: (2026)
von: Li, Zhiyang, et al.
Veröffentlicht: (2026)
Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering
von: Zhang, Zhengxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengxuan, et al.
Veröffentlicht: (2025)
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
von: Lee, Dosung, et al.
Veröffentlicht: (2025)
von: Lee, Dosung, et al.
Veröffentlicht: (2025)
DriveLM: Driving with Graph Visual Question Answering
von: Sima, Chonghao, et al.
Veröffentlicht: (2023)
von: Sima, Chonghao, et al.
Veröffentlicht: (2023)
InViC: Intent-aware Visual Cues for Medical Visual Question Answering
von: Wang, Zhisong, et al.
Veröffentlicht: (2026)
von: Wang, Zhisong, et al.
Veröffentlicht: (2026)
An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
von: Castrillón-Santana, Modesto, et al.
Veröffentlicht: (2025)
von: Castrillón-Santana, Modesto, et al.
Veröffentlicht: (2025)
From Hindsight to Foresight: Self-Encouraged Hindsight Distillation for Knowledge-based Visual Question Answering
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
von: Zhang, Bin, et al.
Veröffentlicht: (2025)
von: Zhang, Bin, et al.
Veröffentlicht: (2025)
See the Forest and the Trees: A Synergistic Reasoning Framework for Knowledge-Based Visual Question Answering
von: Wang, Junjie, et al.
Veröffentlicht: (2025)
von: Wang, Junjie, et al.
Veröffentlicht: (2025)
Knowledge Acquisition Disentanglement for Knowledge-based Visual Question Answering with Large Language Models
von: An, Wenbin, et al.
Veröffentlicht: (2024)
von: An, Wenbin, et al.
Veröffentlicht: (2024)
Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering
von: Hao, Dongze, et al.
Veröffentlicht: (2024)
von: Hao, Dongze, et al.
Veröffentlicht: (2024)
Index Light, Reason Deep: Deferred Visual Ingestion for Visual-Dense Document Question Answering
von: Xu, Tao
Veröffentlicht: (2026)
von: Xu, Tao
Veröffentlicht: (2026)
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
von: Awal, Rabiul, et al.
Veröffentlicht: (2023)
von: Awal, Rabiul, et al.
Veröffentlicht: (2023)
CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering
von: Hong, Yuyang, et al.
Veröffentlicht: (2026)
von: Hong, Yuyang, et al.
Veröffentlicht: (2026)
Detect2Interact: Localizing Object Key Field in Visual Question Answering (VQA) with LLMs
von: Wang, Jialou, et al.
Veröffentlicht: (2024)
von: Wang, Jialou, et al.
Veröffentlicht: (2024)
Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering
von: Chen, Zhuohong, et al.
Veröffentlicht: (2026)
von: Chen, Zhuohong, et al.
Veröffentlicht: (2026)
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
von: Xue, Junxiao, et al.
Veröffentlicht: (2024)
von: Xue, Junxiao, et al.
Veröffentlicht: (2024)
Learning to Compress Contexts for Efficient Knowledge-based Visual Question Answering
von: Weng, Weixi, et al.
Veröffentlicht: (2024)
von: Weng, Weixi, et al.
Veröffentlicht: (2024)
Knowledge Detection by Relevant Question and Image Attributes in Visual Question Answering
von: Ahir, Param, et al.
Veröffentlicht: (2023)
von: Ahir, Param, et al.
Veröffentlicht: (2023)
QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering
von: Zhu, Yingjian, et al.
Veröffentlicht: (2026)
von: Zhu, Yingjian, et al.
Veröffentlicht: (2026)
VQ-VA World: Towards High-Quality Visual Question-Visual Answering
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
V-Loop: Visual Logical Loop Verification for Hallucination Detection in Medical Visual Question Answering
von: Jin, Mengyuan, et al.
Veröffentlicht: (2026)
von: Jin, Mengyuan, et al.
Veröffentlicht: (2026)
Overcoming Language Priors for Visual Question Answering Based on Knowledge Distillation
von: Peng, Daowan, et al.
Veröffentlicht: (2025)
von: Peng, Daowan, et al.
Veröffentlicht: (2025)
Expert Knowledge-Aware Image Difference Graph Representation Learning for Difference-Aware Medical Visual Question Answering
von: Hu, Xinyue, et al.
Veröffentlicht: (2023)
von: Hu, Xinyue, et al.
Veröffentlicht: (2023)
Selectively Answering Visual Questions
von: Eisenschlos, Julian Martin, et al.
Veröffentlicht: (2024)
von: Eisenschlos, Julian Martin, et al.
Veröffentlicht: (2024)
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering
von: Ma, Ziyu, et al.
Veröffentlicht: (2024)
von: Ma, Ziyu, et al.
Veröffentlicht: (2024)
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
von: Cocchi, Federico, et al.
Veröffentlicht: (2024)
von: Cocchi, Federico, et al.
Veröffentlicht: (2024)
Visually Interpretable Subtask Reasoning for Visual Question Answering
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
Targeted Visual Prompting for Medical Visual Question Answering
von: Tascon-Morales, Sergio, et al.
Veröffentlicht: (2024)
von: Tascon-Morales, Sergio, et al.
Veröffentlicht: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
Computed Tomography Visual Question Answering with Cross-modal Feature Graphing
von: Tian, Yuanhe, et al.
Veröffentlicht: (2025)
von: Tian, Yuanhe, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Perceiving Small Visual Details in Zero-shot Visual Question Answering with Multimodal LLMs
von: Zhang, Jiarui, et al.
Veröffentlicht: (2023) -
Knowledge-Enhanced Dual-stream Zero-shot Composed Image Retrieval
von: Suo, Yucheng, et al.
Veröffentlicht: (2024) -
Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation
von: Xu, Quanxing, et al.
Veröffentlicht: (2026) -
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
von: Romero, David, et al.
Veröffentlicht: (2024) -
Map-based Modular Approach for Zero-shot Embodied Question Answering
von: Sakamoto, Koya, et al.
Veröffentlicht: (2024)