Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering
Fuente:
arXiv
Salvato in:
| Autori principali: | Lim, Su Hyeon, Kim, Minkuk, Kim, Hyeon Bae, Kim, Seong Tae |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
di: Kim, Minkuk, et al.
Pubblicazione: (2024)
di: Kim, Minkuk, et al.
Pubblicazione: (2024)
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
di: Kim, Minkuk, et al.
Pubblicazione: (2024)
di: Kim, Minkuk, et al.
Pubblicazione: (2024)
SurgX: Neuron-Concept Association for Explainable Surgical Phase Recognition
di: Kim, Ka Young, et al.
Pubblicazione: (2025)
di: Kim, Ka Young, et al.
Pubblicazione: (2025)
WWW: A Unified Framework for Explaining What, Where and Why of Neural Networks by Interpretation of Neuron Concepts
di: Ahn, Yong Hyun, et al.
Pubblicazione: (2024)
di: Ahn, Yong Hyun, et al.
Pubblicazione: (2024)
Mask-Free Neuron Concept Annotation for Interpreting Neural Networks in Medical Domain
di: Kim, Hyeon Bae, et al.
Pubblicazione: (2024)
di: Kim, Hyeon Bae, et al.
Pubblicazione: (2024)
ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
di: Kim, Ji-Hyeon, et al.
Pubblicazione: (2026)
di: Kim, Ji-Hyeon, et al.
Pubblicazione: (2026)
LLaVA Needs More Knowledge: Retrieval Augmented Natural Language Generation with Knowledge Graph for Explaining Thoracic Pathologies
di: Hamza, Ameer, et al.
Pubblicazione: (2024)
di: Hamza, Ameer, et al.
Pubblicazione: (2024)
Analyzing Effects of Mixed Sample Data Augmentation on Model Interpretability
di: Won, Soyoun, et al.
Pubblicazione: (2023)
di: Won, Soyoun, et al.
Pubblicazione: (2023)
Relevance-aware Multi-context Contrastive Decoding for Retrieval-augmented Visual Question Answering
di: Kim, Jongha, et al.
Pubblicazione: (2026)
di: Kim, Jongha, et al.
Pubblicazione: (2026)
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
di: Lee, Dosung, et al.
Pubblicazione: (2025)
di: Lee, Dosung, et al.
Pubblicazione: (2025)
GradMix: Gradient-based Selective Mixup for Robust Data Augmentation in Class-Incremental Learning
di: Kim, Minsu, et al.
Pubblicazione: (2025)
di: Kim, Minsu, et al.
Pubblicazione: (2025)
UM-Depth : Uncertainty Masked Self-Supervised Monocular Depth Estimation with Visual Odometry
di: Um, Tae-Wook, et al.
Pubblicazione: (2025)
di: Um, Tae-Wook, et al.
Pubblicazione: (2025)
GaussExplorer: 3D Gaussian Splatting for Embodied Exploration and Reasoning
di: Yu-Ji, Kim, et al.
Pubblicazione: (2026)
di: Yu-Ji, Kim, et al.
Pubblicazione: (2026)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
di: Oh, Ju-Young, et al.
Pubblicazione: (2025)
di: Oh, Ju-Young, et al.
Pubblicazione: (2025)
Multimodal Rationales for Explainable Visual Question Answering
di: Li, Kun, et al.
Pubblicazione: (2024)
di: Li, Kun, et al.
Pubblicazione: (2024)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
di: Kim, Hongyeob, et al.
Pubblicazione: (2025)
di: Kim, Hongyeob, et al.
Pubblicazione: (2025)
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering
di: Kil, Jihyung, et al.
Pubblicazione: (2024)
di: Kil, Jihyung, et al.
Pubblicazione: (2024)
Towards Top-Down Reasoning: An Explainable Multi-Agent Approach for Visual Question Answering
di: Wang, Zeqing, et al.
Pubblicazione: (2023)
di: Wang, Zeqing, et al.
Pubblicazione: (2023)
Key-point Guided Deformable Image Manipulation Using Diffusion Model
di: Oh, Seok-Hwan, et al.
Pubblicazione: (2024)
di: Oh, Seok-Hwan, et al.
Pubblicazione: (2024)
Hallucination Benchmark in Medical Visual Question Answering
di: Wu, Jinge, et al.
Pubblicazione: (2024)
di: Wu, Jinge, et al.
Pubblicazione: (2024)
Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation
di: Xu, Quanxing, et al.
Pubblicazione: (2026)
di: Xu, Quanxing, et al.
Pubblicazione: (2026)
Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference
di: Cho, Hyeonwoo, et al.
Pubblicazione: (2026)
di: Cho, Hyeonwoo, et al.
Pubblicazione: (2026)
SEA: Evaluating Sketch Abstraction Efficiency via Element-level Commonsense Visual Question Answering
di: Park, Jiho, et al.
Pubblicazione: (2026)
di: Park, Jiho, et al.
Pubblicazione: (2026)
MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
di: Song, Seokwon, et al.
Pubblicazione: (2025)
di: Song, Seokwon, et al.
Pubblicazione: (2025)
Training-Free Coverless Multi-Image Steganography with Access Control
di: Bae, Minyeol, et al.
Pubblicazione: (2026)
di: Bae, Minyeol, et al.
Pubblicazione: (2026)
Learning Multi-frame and Monocular Prior for Estimating Geometry in Dynamic Scenes
di: Park, Seong Hyeon, et al.
Pubblicazione: (2025)
di: Park, Seong Hyeon, et al.
Pubblicazione: (2025)
Scratching Visual Transformer's Back with Uniform Attention
di: Hyeon-Woo, Nam, et al.
Pubblicazione: (2022)
di: Hyeon-Woo, Nam, et al.
Pubblicazione: (2022)
Empirical Analysis of Anomaly Detection on Hyperspectral Imaging Using Dimension Reduction Methods
di: Kim, Dongeon, et al.
Pubblicazione: (2024)
di: Kim, Dongeon, et al.
Pubblicazione: (2024)
Visually Interpretable Subtask Reasoning for Visual Question Answering
di: Cheng, Yu, et al.
Pubblicazione: (2025)
di: Cheng, Yu, et al.
Pubblicazione: (2025)
Contrastive Language Prompting to Ease False Positives in Medical Anomaly Detection
di: Park, YeongHyeon, et al.
Pubblicazione: (2024)
di: Park, YeongHyeon, et al.
Pubblicazione: (2024)
SnapNTell: Enhancing Entity-Centric Visual Question Answering with Retrieval Augmented Multimodal LLM
di: Qiu, Jielin, et al.
Pubblicazione: (2024)
di: Qiu, Jielin, et al.
Pubblicazione: (2024)
DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning
di: Choi, Joonmyung, et al.
Pubblicazione: (2026)
di: Choi, Joonmyung, et al.
Pubblicazione: (2026)
Pretraining Vision-Language Model for Difference Visual Question Answering in Longitudinal Chest X-rays
di: Cho, Yeongjae, et al.
Pubblicazione: (2024)
di: Cho, Yeongjae, et al.
Pubblicazione: (2024)
PCBEAR: Pose Concept Bottleneck for Explainable Action Recognition
di: Lee, Jongseo, et al.
Pubblicazione: (2025)
di: Lee, Jongseo, et al.
Pubblicazione: (2025)
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
di: Lim, Hyeonseok, et al.
Pubblicazione: (2024)
di: Lim, Hyeonseok, et al.
Pubblicazione: (2024)
Object Retrieval for Visual Question Answering with Outside Knowledge
di: Kan, Shichao, et al.
Pubblicazione: (2024)
di: Kan, Shichao, et al.
Pubblicazione: (2024)
Retrieval-Augmented Open-Vocabulary Object Detection
di: Kim, Jooyeon, et al.
Pubblicazione: (2024)
di: Kim, Jooyeon, et al.
Pubblicazione: (2024)
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning
di: Kim, Geewook, et al.
Pubblicazione: (2024)
di: Kim, Geewook, et al.
Pubblicazione: (2024)
TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation
di: Park, NaHyeon, et al.
Pubblicazione: (2024)
di: Park, NaHyeon, et al.
Pubblicazione: (2024)
RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual Data
di: Cho, Yoorhim, et al.
Pubblicazione: (2025)
di: Cho, Yoorhim, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
di: Kim, Minkuk, et al.
Pubblicazione: (2024) -
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
di: Kim, Minkuk, et al.
Pubblicazione: (2024) -
SurgX: Neuron-Concept Association for Explainable Surgical Phase Recognition
di: Kim, Ka Young, et al.
Pubblicazione: (2025) -
WWW: A Unified Framework for Explaining What, Where and Why of Neural Networks by Interpretation of Neuron Concepts
di: Ahn, Yong Hyun, et al.
Pubblicazione: (2024) -
Mask-Free Neuron Concept Annotation for Interpreting Neural Networks in Medical Domain
di: Kim, Hyeon Bae, et al.
Pubblicazione: (2024)