On the Role of Visual Grounding in VQA
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Reich, Daniel, Schultz, Tanja |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Uncovering the Full Potential of Visual Grounding Methods in VQA
von: Reich, Daniel, et al.
Veröffentlicht: (2024)
von: Reich, Daniel, et al.
Veröffentlicht: (2024)
Measuring Faithful and Plausible Visual Grounding in VQA
von: Reich, Daniel, et al.
Veröffentlicht: (2023)
von: Reich, Daniel, et al.
Veröffentlicht: (2023)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
von: Shrestha, Robik, et al.
Veröffentlicht: (2020)
von: Shrestha, Robik, et al.
Veröffentlicht: (2020)
Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA
von: Safwan, Itbaan, et al.
Veröffentlicht: (2025)
von: Safwan, Itbaan, et al.
Veröffentlicht: (2025)
AutoVQA-G: Self-Improving Agentic Framework for Automated Visual Question Answering and Grounding Annotation
von: Hu, Rongsheng, et al.
Veröffentlicht: (2026)
von: Hu, Rongsheng, et al.
Veröffentlicht: (2026)
Visual Robustness Benchmark for Visual Question Answering (VQA)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
A Modular Pipeline for 3D Object Tracking Using RGB Cameras
von: Bredereke, Lars, et al.
Veröffentlicht: (2025)
von: Bredereke, Lars, et al.
Veröffentlicht: (2025)
CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
von: Qian, Jiahe, et al.
Veröffentlicht: (2025)
von: Qian, Jiahe, et al.
Veröffentlicht: (2025)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
von: Drago, Mauro Orazio, et al.
Veröffentlicht: (2025)
von: Drago, Mauro Orazio, et al.
Veröffentlicht: (2025)
StackOverflowVQA: Stack Overflow Visual Question Answering Dataset
von: Mirzaei, Motahhare, et al.
Veröffentlicht: (2024)
von: Mirzaei, Motahhare, et al.
Veröffentlicht: (2024)
BERT-VQA: Visual Question Answering on Plots
von: Vu, Tai, et al.
Veröffentlicht: (2025)
von: Vu, Tai, et al.
Veröffentlicht: (2025)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
von: Rostamkhani, Mohammadmostafa, et al.
Veröffentlicht: (2024)
von: Rostamkhani, Mohammadmostafa, et al.
Veröffentlicht: (2024)
CommVQA: Situating Visual Question Answering in Communicative Contexts
von: Naik, Nandita Shankar, et al.
Veröffentlicht: (2024)
von: Naik, Nandita Shankar, et al.
Veröffentlicht: (2024)
DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes
von: Al-Mohannadi, Aisha, et al.
Veröffentlicht: (2026)
von: Al-Mohannadi, Aisha, et al.
Veröffentlicht: (2026)
QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA
von: Li, Shuai, et al.
Veröffentlicht: (2025)
von: Li, Shuai, et al.
Veröffentlicht: (2025)
MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
von: Nguyen, Hai-Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Hai-Dang, et al.
Veröffentlicht: (2025)
PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
Head-Aware Visual Cropping: Enhancing Fine-Grained VQA with Attention-Guided Subimage
von: Xie, Junfei, et al.
Veröffentlicht: (2026)
von: Xie, Junfei, et al.
Veröffentlicht: (2026)
Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence
von: Wang, Xingyue, et al.
Veröffentlicht: (2026)
von: Wang, Xingyue, et al.
Veröffentlicht: (2026)
PitVQA: Image-grounded Text Embedding LLM for Visual Question Answering in Pituitary Surgery
von: He, Runlong, et al.
Veröffentlicht: (2024)
von: He, Runlong, et al.
Veröffentlicht: (2024)
Help Me Identify: Is an LLM+VQA System All We Need to Identify Visual Concepts?
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
Detect2Interact: Localizing Object Key Field in Visual Question Answering (VQA) with LLMs
von: Wang, Jialou, et al.
Veröffentlicht: (2024)
von: Wang, Jialou, et al.
Veröffentlicht: (2024)
ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering
von: Tran, Duong T., et al.
Veröffentlicht: (2025)
von: Tran, Duong T., et al.
Veröffentlicht: (2025)
MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question Answering
von: Li, Zhifei, et al.
Veröffentlicht: (2026)
von: Li, Zhifei, et al.
Veröffentlicht: (2026)
RailVQA: A Benchmark and Framework for Efficient Interpretable Visual Cognition in Automatic Train Operation
von: Zhang, Sen, et al.
Veröffentlicht: (2026)
von: Zhang, Sen, et al.
Veröffentlicht: (2026)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
von: Zhang, Chengyi, et al.
Veröffentlicht: (2026)
von: Zhang, Chengyi, et al.
Veröffentlicht: (2026)
VQA$^2$: Visual Question Answering for Video Quality Assessment
von: Jia, Ziheng, et al.
Veröffentlicht: (2024)
von: Jia, Ziheng, et al.
Veröffentlicht: (2024)
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
GPT-4V-AD: Exploring Grounding Potential of VQA-oriented GPT-4V for Zero-shot Anomaly Detection
von: Zhang, Jiangning, et al.
Veröffentlicht: (2023)
von: Zhang, Jiangning, et al.
Veröffentlicht: (2023)
The Role of Entropy in Visual Grounding: Analysis and Optimization
von: Li, Shuo, et al.
Veröffentlicht: (2025)
von: Li, Shuo, et al.
Veröffentlicht: (2025)
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
von: Vu, Sinh Trong, et al.
Veröffentlicht: (2025)
von: Vu, Sinh Trong, et al.
Veröffentlicht: (2025)
VQA-Diff: Exploiting VQA and Diffusion for Zero-Shot Image-to-3D Vehicle Asset Generation in Autonomous Driving
von: Liu, Yibo, et al.
Veröffentlicht: (2024)
von: Liu, Yibo, et al.
Veröffentlicht: (2024)
CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering
von: Hong, Yuyang, et al.
Veröffentlicht: (2026)
von: Hong, Yuyang, et al.
Veröffentlicht: (2026)
VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
Knowledge Condensation and Reasoning for Knowledge-based VQA
von: Hao, Dongze, et al.
Veröffentlicht: (2024)
von: Hao, Dongze, et al.
Veröffentlicht: (2024)
Exploring OCR-augmented Generation for Bilingual VQA
von: Lee, JoonHo, et al.
Veröffentlicht: (2025)
von: Lee, JoonHo, et al.
Veröffentlicht: (2025)
Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
von: Nguyen, Hieu Minh, et al.
Veröffentlicht: (2025)
von: Nguyen, Hieu Minh, et al.
Veröffentlicht: (2025)
LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA
von: Huang, Jing, et al.
Veröffentlicht: (2025)
von: Huang, Jing, et al.
Veröffentlicht: (2025)
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
von: Kahl, Kim-Celine, et al.
Veröffentlicht: (2024)
von: Kahl, Kim-Celine, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Uncovering the Full Potential of Visual Grounding Methods in VQA
von: Reich, Daniel, et al.
Veröffentlicht: (2024) -
Measuring Faithful and Plausible Visual Grounding in VQA
von: Reich, Daniel, et al.
Veröffentlicht: (2023) -
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
von: Shrestha, Robik, et al.
Veröffentlicht: (2020) -
Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA
von: Safwan, Itbaan, et al.
Veröffentlicht: (2025) -
AutoVQA-G: Self-Improving Agentic Framework for Automated Visual Question Answering and Grounding Annotation
von: Hu, Rongsheng, et al.
Veröffentlicht: (2026)