Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Dosung, Jung, Sangwon, Kim, Boyoung, Kim, Minyoung, Kim, Sungyeon, Sung, Junyoung, Seo, Paul Hongsuck |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust Image Self-Recovery against Tampering using Watermark Generation with Pixel Shuffling
by: Kim, Minyoung, et al.
Published: (2025)
by: Kim, Minyoung, et al.
Published: (2025)
ReSCORE: Label-free Iterative Retriever Training for Multi-hop Question Answering with Relevance-Consistency Supervision
by: Lee, Dosung, et al.
Published: (2025)
by: Lee, Dosung, et al.
Published: (2025)
Spectral-Adaptive Modulation Networks for Visual Perception
by: Yun, Guhnoo, et al.
Published: (2025)
by: Yun, Guhnoo, et al.
Published: (2025)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
by: Kim, Hongyeob, et al.
Published: (2025)
by: Kim, Hongyeob, et al.
Published: (2025)
ReTAG: Retrieval-Enhanced, Topic-Augmented Graph-Based Global Sensemaking
by: Kim, Boyoung, et al.
Published: (2025)
by: Kim, Boyoung, et al.
Published: (2025)
SEA: Evaluating Sketch Abstraction Efficiency via Element-level Commonsense Visual Question Answering
by: Park, Jiho, et al.
Published: (2026)
by: Park, Jiho, et al.
Published: (2026)
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering
by: Lim, Su Hyeon, et al.
Published: (2024)
by: Lim, Su Hyeon, et al.
Published: (2024)
Image Diffusion Models Exhibit Emergent Temporal Propagation in Videos
by: Kim, Youngseo, et al.
Published: (2025)
by: Kim, Youngseo, et al.
Published: (2025)
MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
by: Song, Seokwon, et al.
Published: (2025)
by: Song, Seokwon, et al.
Published: (2025)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
by: Park, Kyu Ri, et al.
Published: (2024)
by: Park, Kyu Ri, et al.
Published: (2024)
MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports
by: Kyung, Sunggu, et al.
Published: (2025)
by: Kyung, Sunggu, et al.
Published: (2025)
Direct Diffusion Score Preference Optimization via Stepwise Contrastive Policy-Pair Supervision
by: Kim, Dohyun, et al.
Published: (2025)
by: Kim, Dohyun, et al.
Published: (2025)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
by: Choi, Changin, et al.
Published: (2025)
by: Choi, Changin, et al.
Published: (2025)
Hallucination Benchmark in Medical Visual Question Answering
by: Wu, Jinge, et al.
Published: (2024)
by: Wu, Jinge, et al.
Published: (2024)
Learning Correlation Structures for Vision Transformers
by: Kim, Manjin, et al.
Published: (2024)
by: Kim, Manjin, et al.
Published: (2024)
Uncertainty-aware Semantic Mapping in Off-road Environments with Dempster-Shafer Theory of Evidence
by: Kim, Junyoung, et al.
Published: (2024)
by: Kim, Junyoung, et al.
Published: (2024)
Object-aware Sound Source Localization via Audio-Visual Scene Understanding
by: Um, Sung Jin, et al.
Published: (2025)
by: Um, Sung Jin, et al.
Published: (2025)
FREST: Feature RESToration for Semantic Segmentation under Multiple Adverse Conditions
by: Lee, Sohyun, et al.
Published: (2024)
by: Lee, Sohyun, et al.
Published: (2024)
QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering
by: Jung, Woojun, et al.
Published: (2026)
by: Jung, Woojun, et al.
Published: (2026)
StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback
by: Park, Jiho, et al.
Published: (2025)
by: Park, Jiho, et al.
Published: (2025)
Multimodal Rationales for Explainable Visual Question Answering
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
Multi-Granularity Video Object Segmentation
by: Lim, Sangbeom, et al.
Published: (2024)
by: Lim, Sangbeom, et al.
Published: (2024)
MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated Label
by: Jung, Junyoung, et al.
Published: (2026)
by: Jung, Junyoung, et al.
Published: (2026)
Relevance-aware Multi-context Contrastive Decoding for Retrieval-augmented Visual Question Answering
by: Kim, Jongha, et al.
Published: (2026)
by: Kim, Jongha, et al.
Published: (2026)
Visual Question Answering Instruction: Unlocking Multimodal Large Language Model To Domain-Specific Visual Multitasks
by: Lee, Jusung, et al.
Published: (2024)
by: Lee, Jusung, et al.
Published: (2024)
Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge
by: Kim, Dongjin, et al.
Published: (2024)
by: Kim, Dongjin, et al.
Published: (2024)
Object Retrieval for Visual Question Answering with Outside Knowledge
by: Kan, Shichao, et al.
Published: (2024)
by: Kan, Shichao, et al.
Published: (2024)
Overcoming Language Priors for Visual Question Answering Based on Knowledge Distillation
by: Peng, Daowan, et al.
Published: (2025)
by: Peng, Daowan, et al.
Published: (2025)
Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering
by: Chen, Zhuohong, et al.
Published: (2026)
by: Chen, Zhuohong, et al.
Published: (2026)
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
by: Xue, Junxiao, et al.
Published: (2024)
by: Xue, Junxiao, et al.
Published: (2024)
MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering
by: Peng, Jingwei, et al.
Published: (2025)
by: Peng, Jingwei, et al.
Published: (2025)
A Visual Question Answering Method for SAR Ship: Breaking the Requirement for Multimodal Dataset Construction and Model Fine-Tuning
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning
by: Choi, Joonmyung, et al.
Published: (2026)
by: Choi, Joonmyung, et al.
Published: (2026)
Visual Representation Alignment for Multimodal Large Language Models
by: Yoon, Heeji, et al.
Published: (2025)
by: Yoon, Heeji, et al.
Published: (2025)
Crafting Query-Aware Selective Attention for Single Image Super-Resolution
by: Kim, Junyoung, et al.
Published: (2025)
by: Kim, Junyoung, et al.
Published: (2025)
WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
by: Yu, Seungjun, et al.
Published: (2025)
by: Yu, Seungjun, et al.
Published: (2025)
Towards Open-Vocabulary Semantic Segmentation Without Semantic Labels
by: Shin, Heeseong, et al.
Published: (2024)
by: Shin, Heeseong, et al.
Published: (2024)
TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains
by: Kim, Yoonsik, et al.
Published: (2024)
by: Kim, Yoonsik, et al.
Published: (2024)
Evidential Semantic Mapping in Off-road Environments with Uncertainty-aware Bayesian Kernel Inference
by: Kim, Junyoung, et al.
Published: (2024)
by: Kim, Junyoung, et al.
Published: (2024)
MedCFVQA: A Causal Approach to Mitigate Modality Preference Bias in Medical Visual Question Answering
by: Ye, Shuchang, et al.
Published: (2025)
by: Ye, Shuchang, et al.
Published: (2025)
Similar Items
-
Robust Image Self-Recovery against Tampering using Watermark Generation with Pixel Shuffling
by: Kim, Minyoung, et al.
Published: (2025) -
ReSCORE: Label-free Iterative Retriever Training for Multi-hop Question Answering with Relevance-Consistency Supervision
by: Lee, Dosung, et al.
Published: (2025) -
Spectral-Adaptive Modulation Networks for Visual Perception
by: Yun, Guhnoo, et al.
Published: (2025) -
Question-Aware Gaussian Experts for Audio-Visual Question Answering
by: Kim, Hongyeob, et al.
Published: (2025) -
ReTAG: Retrieval-Enhanced, Topic-Augmented Graph-Based Global Sensemaking
by: Kim, Boyoung, et al.
Published: (2025)