Why context matters in VQA and Reasoning: Semantic interventions for VLM input modalities
Fuente:
arXiv
Saved in:
| Main Authors: | Amara, Kenza, Klein, Lukas, Lüth, Carsten, Jäger, Paul, Strobelt, Hendrik, El-Assady, Mennatallah |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SyntaxShap: Syntax-aware Explainability Method for Text Generation
by: Amara, Kenza, et al.
Published: (2024)
by: Amara, Kenza, et al.
Published: (2024)
Challenges and Opportunities in Text Generation Explainability
by: Amara, Kenza, et al.
Published: (2024)
by: Amara, Kenza, et al.
Published: (2024)
Concept-Level Explainability for Auditing & Steering LLM Responses
by: Amara, Kenza, et al.
Published: (2025)
by: Amara, Kenza, et al.
Published: (2025)
Deconstructing Human-AI Collaboration: Agency, Interaction, and Adaptation
by: Holter, Steffen, et al.
Published: (2024)
by: Holter, Steffen, et al.
Published: (2024)
Navigating the Maze of Explainable AI: A Systematic Approach to Evaluating Methods and Metrics
by: Klein, Lukas, et al.
Published: (2024)
by: Klein, Lukas, et al.
Published: (2024)
Understanding Large Language Model Behaviors through Interactive Counterfactual Generation and Analysis
by: Cheng, Furui, et al.
Published: (2024)
by: Cheng, Furui, et al.
Published: (2024)
Cross-Cultural Simulation of Citizen Emotional Responses to Bureaucratic Red Tape Using LLM Agents
by: Ni, Wanchun, et al.
Published: (2026)
by: Ni, Wanchun, et al.
Published: (2026)
Reward Learning from Multiple Feedback Types
by: Metz, Yannick, et al.
Published: (2025)
by: Metz, Yannick, et al.
Published: (2025)
PleaSQLarify: Visual Pragmatic Repair for Natural Language Database Querying
by: Chan, Robin Shing Moon, et al.
Published: (2026)
by: Chan, Robin Shing Moon, et al.
Published: (2026)
PowerGraph: A power grid benchmark dataset for graph neural networks
by: Varbella, Anna, et al.
Published: (2024)
by: Varbella, Anna, et al.
Published: (2024)
TopoAlign: Topology-Aware Visual Representation Alignment
by: Yan, Xinyuan, et al.
Published: (2026)
by: Yan, Xinyuan, et al.
Published: (2026)
DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition
by: Shi, Danqing, et al.
Published: (2025)
by: Shi, Danqing, et al.
Published: (2025)
iToT: An Interactive System for Customized Tree-of-Thought Generation
by: Boyle, Alan, et al.
Published: (2024)
by: Boyle, Alan, et al.
Published: (2024)
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
by: Baur, Raphaël, et al.
Published: (2026)
by: Baur, Raphaël, et al.
Published: (2026)
RELIC: Investigating Large Language Model Responses using Self-Consistency
by: Cheng, Furui, et al.
Published: (2023)
by: Cheng, Furui, et al.
Published: (2023)
GPT-2 Through the Lens of Vector Symbolic Architectures
by: Knittel, Johannes, et al.
Published: (2024)
by: Knittel, Johannes, et al.
Published: (2024)
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
by: Bousselham, Walid, et al.
Published: (2026)
by: Bousselham, Walid, et al.
Published: (2026)
ToolVQA: A Dataset for Multi-step Reasoning VQA with External Tools
by: Yin, Shaofeng, et al.
Published: (2025)
by: Yin, Shaofeng, et al.
Published: (2025)
Deconstructing Human‐AI Collaboration: Agency, Interaction, and Adaptation
by: Steffen Holter, et al.
Published: (2024)
by: Steffen Holter, et al.
Published: (2024)
Multi-modal Reasoning with LLMs for Visual Semantic Arithmetic
by: Xu, Chuou, et al.
Published: (2026)
by: Xu, Chuou, et al.
Published: (2026)
Abstraction Alignment: Comparing Model-Learned and Human-Encoded Conceptual Relationships
by: Boggust, Angie, et al.
Published: (2024)
by: Boggust, Angie, et al.
Published: (2024)
Position: Restructuring of Categories and Implementation of Guidelines Essential for VLM Adoption in Healthcare
by: Tariq, Amara, et al.
Published: (2025)
by: Tariq, Amara, et al.
Published: (2025)
Disentangling Knowledge-based and Visual Reasoning by Question Decomposition in KB-VQA
by: Barezi, Elham J., et al.
Published: (2024)
by: Barezi, Elham J., et al.
Published: (2024)
FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis
by: Wong, Zhen Hao, et al.
Published: (2025)
by: Wong, Zhen Hao, et al.
Published: (2025)
Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge
by: Xu, Yinsong, et al.
Published: (2026)
by: Xu, Yinsong, et al.
Published: (2026)
Architecting Clinical Collaboration: Multi-Agent Reasoning Systems for Multimodal Medical VQA
by: Thakrar, Karishma, et al.
Published: (2025)
by: Thakrar, Karishma, et al.
Published: (2025)
KEPO: Knowledge-Enhanced Preference Optimization for Multimodal Reasoning with Applications to Medical VQA
by: Yang, Fan, et al.
Published: (2026)
by: Yang, Fan, et al.
Published: (2026)
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
by: Paglieri, Davide, et al.
Published: (2024)
by: Paglieri, Davide, et al.
Published: (2024)
Memory in Plain Sight: Surveying the Uncanny Resemblances of Associative Memories and Diffusion Models
by: Hoover, Benjamin, et al.
Published: (2023)
by: Hoover, Benjamin, et al.
Published: (2023)
iNNspector: Visual, Interactive Deep Model Debugging
by: Spinner, Thilo, et al.
Published: (2024)
by: Spinner, Thilo, et al.
Published: (2024)
GraphFramEx: Towards Systematic Evaluation of Explainability Methods for Graph Neural Networks
by: Amara, Kenza, et al.
Published: (2022)
by: Amara, Kenza, et al.
Published: (2022)
Cross-Stage Coherence in Hierarchical Driving VQA: Explicit Baselines and Learned Gated Context Projectors
by: Jain, Gautam Kumar, et al.
Published: (2026)
by: Jain, Gautam Kumar, et al.
Published: (2026)
A Semantic Autonomy Framework for VLM-Integrated Indoor Mobile Robots: Hybrid Deterministic Reasoning and Cross-Robot Adaptive Memory
by: Abaza, Bogdan Felician, et al.
Published: (2026)
by: Abaza, Bogdan Felician, et al.
Published: (2026)
Small Models, Smarter Learning: The Power of Joint Task Training
by: Both, Csaba, et al.
Published: (2025)
by: Both, Csaba, et al.
Published: (2025)
Simulated Reasoning is Reasoning
by: Kempt, Hendrik, et al.
Published: (2026)
by: Kempt, Hendrik, et al.
Published: (2026)
R^3-VQA: "Read the Room" by Video Social Reasoning
by: Niu, Lixing, et al.
Published: (2025)
by: Niu, Lixing, et al.
Published: (2025)
SafeCoT: Improving VLM Safety with Minimal Reasoning
by: Ma, Jiachen, et al.
Published: (2025)
by: Ma, Jiachen, et al.
Published: (2025)
MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents
by: Zhang, Ruoxuan, et al.
Published: (2025)
by: Zhang, Ruoxuan, et al.
Published: (2025)
SD-E$^2$: Semantic Exploration for Reasoning Under Token Budgets
by: Mishra, Kshitij, et al.
Published: (2026)
by: Mishra, Kshitij, et al.
Published: (2026)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
by: Shrestha, Robik, et al.
Published: (2020)
by: Shrestha, Robik, et al.
Published: (2020)
Similar Items
-
SyntaxShap: Syntax-aware Explainability Method for Text Generation
by: Amara, Kenza, et al.
Published: (2024) -
Challenges and Opportunities in Text Generation Explainability
by: Amara, Kenza, et al.
Published: (2024) -
Concept-Level Explainability for Auditing & Steering LLM Responses
by: Amara, Kenza, et al.
Published: (2025) -
Deconstructing Human-AI Collaboration: Agency, Interaction, and Adaptation
by: Holter, Steffen, et al.
Published: (2024) -
Navigating the Maze of Explainable AI: A Systematic Approach to Evaluating Methods and Metrics
by: Klein, Lukas, et al.
Published: (2024)