MaS-VQA: A Mask-and-Select Framework for Knowledge-Based Visual Question Answering
Fuente:
arXiv
Guardado en:
| Autores principales: | Mao, Xianwei, Ye, Kai, Zhou, Sheng, Zhang, Nan, Huang, Haikuan, Li, Bin, Bu, Jiajun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot Alignment
por: Ye, Kai, et al.
Publicado: (2026)
por: Ye, Kai, et al.
Publicado: (2026)
CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering
por: Hong, Yuyang, et al.
Publicado: (2026)
por: Hong, Yuyang, et al.
Publicado: (2026)
Visual Robustness Benchmark for Visual Question Answering (VQA)
por: Ishmam, Md Farhan, et al.
Publicado: (2024)
por: Ishmam, Md Farhan, et al.
Publicado: (2024)
BERT-VQA: Visual Question Answering on Plots
por: Vu, Tai, et al.
Publicado: (2025)
por: Vu, Tai, et al.
Publicado: (2025)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
por: Zhang, Xiaoman, et al.
Publicado: (2023)
por: Zhang, Xiaoman, et al.
Publicado: (2023)
ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering
por: Tran, Duong T., et al.
Publicado: (2025)
por: Tran, Duong T., et al.
Publicado: (2025)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
por: Zhang, Chengyi, et al.
Publicado: (2026)
por: Zhang, Chengyi, et al.
Publicado: (2026)
StackOverflowVQA: Stack Overflow Visual Question Answering Dataset
por: Mirzaei, Motahhare, et al.
Publicado: (2024)
por: Mirzaei, Motahhare, et al.
Publicado: (2024)
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
por: Mo, Ye, et al.
Publicado: (2025)
por: Mo, Ye, et al.
Publicado: (2025)
CommVQA: Situating Visual Question Answering in Communicative Contexts
por: Naik, Nandita Shankar, et al.
Publicado: (2024)
por: Naik, Nandita Shankar, et al.
Publicado: (2024)
VQA$^2$: Visual Question Answering for Video Quality Assessment
por: Jia, Ziheng, et al.
Publicado: (2024)
por: Jia, Ziheng, et al.
Publicado: (2024)
See the Forest and the Trees: A Synergistic Reasoning Framework for Knowledge-Based Visual Question Answering
por: Wang, Junjie, et al.
Publicado: (2025)
por: Wang, Junjie, et al.
Publicado: (2025)
AutoVQA-G: Self-Improving Agentic Framework for Automated Visual Question Answering and Grounding Annotation
por: Hu, Rongsheng, et al.
Publicado: (2026)
por: Hu, Rongsheng, et al.
Publicado: (2026)
DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes
por: Al-Mohannadi, Aisha, et al.
Publicado: (2026)
por: Al-Mohannadi, Aisha, et al.
Publicado: (2026)
MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
por: Nguyen, Hai-Dang, et al.
Publicado: (2025)
por: Nguyen, Hai-Dang, et al.
Publicado: (2025)
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
por: Vu, Sinh Trong, et al.
Publicado: (2025)
por: Vu, Sinh Trong, et al.
Publicado: (2025)
ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering
por: Diao, Xingjian, et al.
Publicado: (2025)
por: Diao, Xingjian, et al.
Publicado: (2025)
Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering
por: Zhang, Zhengxuan, et al.
Publicado: (2025)
por: Zhang, Zhengxuan, et al.
Publicado: (2025)
Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering
por: Hao, Dongze, et al.
Publicado: (2024)
por: Hao, Dongze, et al.
Publicado: (2024)
WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
por: Chen, Pingyi, et al.
Publicado: (2024)
por: Chen, Pingyi, et al.
Publicado: (2024)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
por: Zhou, Sheng, et al.
Publicado: (2025)
por: Zhou, Sheng, et al.
Publicado: (2025)
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
por: Lee, Dosung, et al.
Publicado: (2025)
por: Lee, Dosung, et al.
Publicado: (2025)
Selectively Answering Visual Questions
por: Eisenschlos, Julian Martin, et al.
Publicado: (2024)
por: Eisenschlos, Julian Martin, et al.
Publicado: (2024)
PitVQA: Image-grounded Text Embedding LLM for Visual Question Answering in Pituitary Surgery
por: He, Runlong, et al.
Publicado: (2024)
por: He, Runlong, et al.
Publicado: (2024)
MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question Answering
por: Li, Zhifei, et al.
Publicado: (2026)
por: Li, Zhifei, et al.
Publicado: (2026)
Detect2Interact: Localizing Object Key Field in Visual Question Answering (VQA) with LLMs
por: Wang, Jialou, et al.
Publicado: (2024)
por: Wang, Jialou, et al.
Publicado: (2024)
Reconstruction as a Bridge for Event-Based Visual Question Answering
por: Lou, Hanyue, et al.
Publicado: (2025)
por: Lou, Hanyue, et al.
Publicado: (2025)
TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains
por: Kim, Yoonsik, et al.
Publicado: (2024)
por: Kim, Yoonsik, et al.
Publicado: (2024)
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts
por: Singh, Shubhankar, et al.
Publicado: (2024)
por: Singh, Shubhankar, et al.
Publicado: (2024)
Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
por: Nguyen, Hieu Minh, et al.
Publicado: (2025)
por: Nguyen, Hieu Minh, et al.
Publicado: (2025)
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
por: Guan, Runwei, et al.
Publicado: (2025)
por: Guan, Runwei, et al.
Publicado: (2025)
A Knowledge Noise Mitigation Framework for Knowledge-based Visual Question Answering
por: Liu, Zhiyue, et al.
Publicado: (2025)
por: Liu, Zhiyue, et al.
Publicado: (2025)
MGA-VQA: Secure and Interpretable Graph-Augmented Visual Question Answering with Memory-Guided Protection Against Unauthorized Knowledge Use
por: Mohammadshirazi, Ahmad, et al.
Publicado: (2025)
por: Mohammadshirazi, Ahmad, et al.
Publicado: (2025)
VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering
por: Sood, Ekta, et al.
Publicado: (2021)
por: Sood, Ekta, et al.
Publicado: (2021)
SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering
por: Zhang, Yan, et al.
Publicado: (2025)
por: Zhang, Yan, et al.
Publicado: (2025)
Adversarial Attacks on VQA-NLE: Exposing and Alleviating Inconsistencies in Visual Question Answering Explanations
por: Yeh, Yahsin, et al.
Publicado: (2025)
por: Yeh, Yahsin, et al.
Publicado: (2025)
UCSF-PDGM-VQA: Visual Question Answering dataset for brain tumor MRI interpretation
por: Ghosh, Shiv, et al.
Publicado: (2026)
por: Ghosh, Shiv, et al.
Publicado: (2026)
Object Retrieval for Visual Question Answering with Outside Knowledge
por: Kan, Shichao, et al.
Publicado: (2024)
por: Kan, Shichao, et al.
Publicado: (2024)
Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering
por: Chen, Zhuohong, et al.
Publicado: (2026)
por: Chen, Zhuohong, et al.
Publicado: (2026)
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering
por: Ma, Ziyu, et al.
Publicado: (2024)
por: Ma, Ziyu, et al.
Publicado: (2024)
Ejemplares similares
-
REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot Alignment
por: Ye, Kai, et al.
Publicado: (2026) -
CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering
por: Hong, Yuyang, et al.
Publicado: (2026) -
Visual Robustness Benchmark for Visual Question Answering (VQA)
por: Ishmam, Md Farhan, et al.
Publicado: (2024) -
BERT-VQA: Visual Question Answering on Plots
por: Vu, Tai, et al.
Publicado: (2025) -
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
por: Zhang, Xiaoman, et al.
Publicado: (2023)