Measuring Faithful and Plausible Visual Grounding in VQA
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Reich, Daniel, Putze, Felix, Schultz, Tanja |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Role of Visual Grounding in VQA
von: Reich, Daniel, et al.
Veröffentlicht: (2024)
von: Reich, Daniel, et al.
Veröffentlicht: (2024)
Uncovering the Full Potential of Visual Grounding Methods in VQA
von: Reich, Daniel, et al.
Veröffentlicht: (2024)
von: Reich, Daniel, et al.
Veröffentlicht: (2024)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
von: Shrestha, Robik, et al.
Veröffentlicht: (2020)
von: Shrestha, Robik, et al.
Veröffentlicht: (2020)
Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA
von: Safwan, Itbaan, et al.
Veröffentlicht: (2025)
von: Safwan, Itbaan, et al.
Veröffentlicht: (2025)
On the Faithfulness of Visual Thinking: Measurement and Enhancement
von: Liu, Zujing, et al.
Veröffentlicht: (2025)
von: Liu, Zujing, et al.
Veröffentlicht: (2025)
AutoVQA-G: Self-Improving Agentic Framework for Automated Visual Question Answering and Grounding Annotation
von: Hu, Rongsheng, et al.
Veröffentlicht: (2026)
von: Hu, Rongsheng, et al.
Veröffentlicht: (2026)
Visual Robustness Benchmark for Visual Question Answering (VQA)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
A Modular Pipeline for 3D Object Tracking Using RGB Cameras
von: Bredereke, Lars, et al.
Veröffentlicht: (2025)
von: Bredereke, Lars, et al.
Veröffentlicht: (2025)
CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
von: Qian, Jiahe, et al.
Veröffentlicht: (2025)
von: Qian, Jiahe, et al.
Veröffentlicht: (2025)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
Faithful Counterfactual Visual Explanations (FCVE)
von: Khan, Bismillah, et al.
Veröffentlicht: (2025)
von: Khan, Bismillah, et al.
Veröffentlicht: (2025)
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
von: Drago, Mauro Orazio, et al.
Veröffentlicht: (2025)
von: Drago, Mauro Orazio, et al.
Veröffentlicht: (2025)
StackOverflowVQA: Stack Overflow Visual Question Answering Dataset
von: Mirzaei, Motahhare, et al.
Veröffentlicht: (2024)
von: Mirzaei, Motahhare, et al.
Veröffentlicht: (2024)
Measuring Physical Plausibility of 3D Human Poses Using Physics Simulation
von: Louis, Nathan, et al.
Veröffentlicht: (2025)
von: Louis, Nathan, et al.
Veröffentlicht: (2025)
BERT-VQA: Visual Question Answering on Plots
von: Vu, Tai, et al.
Veröffentlicht: (2025)
von: Vu, Tai, et al.
Veröffentlicht: (2025)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
von: Rostamkhani, Mohammadmostafa, et al.
Veröffentlicht: (2024)
von: Rostamkhani, Mohammadmostafa, et al.
Veröffentlicht: (2024)
CommVQA: Situating Visual Question Answering in Communicative Contexts
von: Naik, Nandita Shankar, et al.
Veröffentlicht: (2024)
von: Naik, Nandita Shankar, et al.
Veröffentlicht: (2024)
DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes
von: Al-Mohannadi, Aisha, et al.
Veröffentlicht: (2026)
von: Al-Mohannadi, Aisha, et al.
Veröffentlicht: (2026)
QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA
von: Li, Shuai, et al.
Veröffentlicht: (2025)
von: Li, Shuai, et al.
Veröffentlicht: (2025)
MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
von: Nguyen, Hai-Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Hai-Dang, et al.
Veröffentlicht: (2025)
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
von: Tian, Changyuan, et al.
Veröffentlicht: (2026)
von: Tian, Changyuan, et al.
Veröffentlicht: (2026)
Head-Aware Visual Cropping: Enhancing Fine-Grained VQA with Attention-Guided Subimage
von: Xie, Junfei, et al.
Veröffentlicht: (2026)
von: Xie, Junfei, et al.
Veröffentlicht: (2026)
PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence
von: Wang, Xingyue, et al.
Veröffentlicht: (2026)
von: Wang, Xingyue, et al.
Veröffentlicht: (2026)
ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering
von: Tran, Duong T., et al.
Veröffentlicht: (2025)
von: Tran, Duong T., et al.
Veröffentlicht: (2025)
PitVQA: Image-grounded Text Embedding LLM for Visual Question Answering in Pituitary Surgery
von: He, Runlong, et al.
Veröffentlicht: (2024)
von: He, Runlong, et al.
Veröffentlicht: (2024)
MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question Answering
von: Li, Zhifei, et al.
Veröffentlicht: (2026)
von: Li, Zhifei, et al.
Veröffentlicht: (2026)
Help Me Identify: Is an LLM+VQA System All We Need to Identify Visual Concepts?
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
Detect2Interact: Localizing Object Key Field in Visual Question Answering (VQA) with LLMs
von: Wang, Jialou, et al.
Veröffentlicht: (2024)
von: Wang, Jialou, et al.
Veröffentlicht: (2024)
RailVQA: A Benchmark and Framework for Efficient Interpretable Visual Cognition in Automatic Train Operation
von: Zhang, Sen, et al.
Veröffentlicht: (2026)
von: Zhang, Sen, et al.
Veröffentlicht: (2026)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
von: Zhang, Chengyi, et al.
Veröffentlicht: (2026)
von: Zhang, Chengyi, et al.
Veröffentlicht: (2026)
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
von: Rahman, Md Ashikur, et al.
Veröffentlicht: (2026)
von: Rahman, Md Ashikur, et al.
Veröffentlicht: (2026)
Visual Tomography: Physically Faithful Volumetric Models of Partially Translucent Objects
von: Nakath, David, et al.
Veröffentlicht: (2023)
von: Nakath, David, et al.
Veröffentlicht: (2023)
VQA$^2$: Visual Question Answering for Video Quality Assessment
von: Jia, Ziheng, et al.
Veröffentlicht: (2024)
von: Jia, Ziheng, et al.
Veröffentlicht: (2024)
As-Plausible-As-Possible: Plausibility-Aware Mesh Deformation Using 2D Diffusion Priors
von: Yoo, Seungwoo, et al.
Veröffentlicht: (2023)
von: Yoo, Seungwoo, et al.
Veröffentlicht: (2023)
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
GPT-4V-AD: Exploring Grounding Potential of VQA-oriented GPT-4V for Zero-shot Anomaly Detection
von: Zhang, Jiangning, et al.
Veröffentlicht: (2023)
von: Zhang, Jiangning, et al.
Veröffentlicht: (2023)
Motion Diffusion Autoencoders: Enabling Attribute Manipulation in Human Motion Demonstrated on Karate Techniques
von: Richardson, Anthony, et al.
Veröffentlicht: (2025)
von: Richardson, Anthony, et al.
Veröffentlicht: (2025)
Improving Human Motion Plausibility with Body Momentum
von: Nguyen, Ha Linh, et al.
Veröffentlicht: (2025)
von: Nguyen, Ha Linh, et al.
Veröffentlicht: (2025)
EditTransfer++: Toward Faithful and Efficient Visual-Prompt-Guided Image Editing
von: Chen, Lan, et al.
Veröffentlicht: (2026)
von: Chen, Lan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
On the Role of Visual Grounding in VQA
von: Reich, Daniel, et al.
Veröffentlicht: (2024) -
Uncovering the Full Potential of Visual Grounding Methods in VQA
von: Reich, Daniel, et al.
Veröffentlicht: (2024) -
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
von: Shrestha, Robik, et al.
Veröffentlicht: (2020) -
Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA
von: Safwan, Itbaan, et al.
Veröffentlicht: (2025) -
On the Faithfulness of Visual Thinking: Measurement and Enhancement
von: Liu, Zujing, et al.
Veröffentlicht: (2025)