Why context matters in VQA and Reasoning: Semantic interventions for VLM input modalities
Fuente:
arXiv
Guardado en:
| Autores principales: | Amara, Kenza, Klein, Lukas, Lüth, Carsten, Jäger, Paul, Strobelt, Hendrik, El-Assady, Mennatallah |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SyntaxShap: Syntax-aware Explainability Method for Text Generation
por: Amara, Kenza, et al.
Publicado: (2024)
por: Amara, Kenza, et al.
Publicado: (2024)
Challenges and Opportunities in Text Generation Explainability
por: Amara, Kenza, et al.
Publicado: (2024)
por: Amara, Kenza, et al.
Publicado: (2024)
Concept-Level Explainability for Auditing & Steering LLM Responses
por: Amara, Kenza, et al.
Publicado: (2025)
por: Amara, Kenza, et al.
Publicado: (2025)
Deconstructing Human-AI Collaboration: Agency, Interaction, and Adaptation
por: Holter, Steffen, et al.
Publicado: (2024)
por: Holter, Steffen, et al.
Publicado: (2024)
Navigating the Maze of Explainable AI: A Systematic Approach to Evaluating Methods and Metrics
por: Klein, Lukas, et al.
Publicado: (2024)
por: Klein, Lukas, et al.
Publicado: (2024)
Understanding Large Language Model Behaviors through Interactive Counterfactual Generation and Analysis
por: Cheng, Furui, et al.
Publicado: (2024)
por: Cheng, Furui, et al.
Publicado: (2024)
Cross-Cultural Simulation of Citizen Emotional Responses to Bureaucratic Red Tape Using LLM Agents
por: Ni, Wanchun, et al.
Publicado: (2026)
por: Ni, Wanchun, et al.
Publicado: (2026)
Reward Learning from Multiple Feedback Types
por: Metz, Yannick, et al.
Publicado: (2025)
por: Metz, Yannick, et al.
Publicado: (2025)
PleaSQLarify: Visual Pragmatic Repair for Natural Language Database Querying
por: Chan, Robin Shing Moon, et al.
Publicado: (2026)
por: Chan, Robin Shing Moon, et al.
Publicado: (2026)
PowerGraph: A power grid benchmark dataset for graph neural networks
por: Varbella, Anna, et al.
Publicado: (2024)
por: Varbella, Anna, et al.
Publicado: (2024)
TopoAlign: Topology-Aware Visual Representation Alignment
por: Yan, Xinyuan, et al.
Publicado: (2026)
por: Yan, Xinyuan, et al.
Publicado: (2026)
DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition
por: Shi, Danqing, et al.
Publicado: (2025)
por: Shi, Danqing, et al.
Publicado: (2025)
iToT: An Interactive System for Customized Tree-of-Thought Generation
por: Boyle, Alan, et al.
Publicado: (2024)
por: Boyle, Alan, et al.
Publicado: (2024)
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
por: Baur, Raphaël, et al.
Publicado: (2026)
por: Baur, Raphaël, et al.
Publicado: (2026)
RELIC: Investigating Large Language Model Responses using Self-Consistency
por: Cheng, Furui, et al.
Publicado: (2023)
por: Cheng, Furui, et al.
Publicado: (2023)
GPT-2 Through the Lens of Vector Symbolic Architectures
por: Knittel, Johannes, et al.
Publicado: (2024)
por: Knittel, Johannes, et al.
Publicado: (2024)
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
por: Bousselham, Walid, et al.
Publicado: (2026)
por: Bousselham, Walid, et al.
Publicado: (2026)
ToolVQA: A Dataset for Multi-step Reasoning VQA with External Tools
por: Yin, Shaofeng, et al.
Publicado: (2025)
por: Yin, Shaofeng, et al.
Publicado: (2025)
Deconstructing Human‐AI Collaboration: Agency, Interaction, and Adaptation
por: Steffen Holter, et al.
Publicado: (2024)
por: Steffen Holter, et al.
Publicado: (2024)
Multi-modal Reasoning with LLMs for Visual Semantic Arithmetic
por: Xu, Chuou, et al.
Publicado: (2026)
por: Xu, Chuou, et al.
Publicado: (2026)
Abstraction Alignment: Comparing Model-Learned and Human-Encoded Conceptual Relationships
por: Boggust, Angie, et al.
Publicado: (2024)
por: Boggust, Angie, et al.
Publicado: (2024)
Position: Restructuring of Categories and Implementation of Guidelines Essential for VLM Adoption in Healthcare
por: Tariq, Amara, et al.
Publicado: (2025)
por: Tariq, Amara, et al.
Publicado: (2025)
Disentangling Knowledge-based and Visual Reasoning by Question Decomposition in KB-VQA
por: Barezi, Elham J., et al.
Publicado: (2024)
por: Barezi, Elham J., et al.
Publicado: (2024)
FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis
por: Wong, Zhen Hao, et al.
Publicado: (2025)
por: Wong, Zhen Hao, et al.
Publicado: (2025)
Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge
por: Xu, Yinsong, et al.
Publicado: (2026)
por: Xu, Yinsong, et al.
Publicado: (2026)
Architecting Clinical Collaboration: Multi-Agent Reasoning Systems for Multimodal Medical VQA
por: Thakrar, Karishma, et al.
Publicado: (2025)
por: Thakrar, Karishma, et al.
Publicado: (2025)
KEPO: Knowledge-Enhanced Preference Optimization for Multimodal Reasoning with Applications to Medical VQA
por: Yang, Fan, et al.
Publicado: (2026)
por: Yang, Fan, et al.
Publicado: (2026)
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
por: Paglieri, Davide, et al.
Publicado: (2024)
por: Paglieri, Davide, et al.
Publicado: (2024)
Memory in Plain Sight: Surveying the Uncanny Resemblances of Associative Memories and Diffusion Models
por: Hoover, Benjamin, et al.
Publicado: (2023)
por: Hoover, Benjamin, et al.
Publicado: (2023)
iNNspector: Visual, Interactive Deep Model Debugging
por: Spinner, Thilo, et al.
Publicado: (2024)
por: Spinner, Thilo, et al.
Publicado: (2024)
GraphFramEx: Towards Systematic Evaluation of Explainability Methods for Graph Neural Networks
por: Amara, Kenza, et al.
Publicado: (2022)
por: Amara, Kenza, et al.
Publicado: (2022)
Cross-Stage Coherence in Hierarchical Driving VQA: Explicit Baselines and Learned Gated Context Projectors
por: Jain, Gautam Kumar, et al.
Publicado: (2026)
por: Jain, Gautam Kumar, et al.
Publicado: (2026)
A Semantic Autonomy Framework for VLM-Integrated Indoor Mobile Robots: Hybrid Deterministic Reasoning and Cross-Robot Adaptive Memory
por: Abaza, Bogdan Felician, et al.
Publicado: (2026)
por: Abaza, Bogdan Felician, et al.
Publicado: (2026)
Small Models, Smarter Learning: The Power of Joint Task Training
por: Both, Csaba, et al.
Publicado: (2025)
por: Both, Csaba, et al.
Publicado: (2025)
Simulated Reasoning is Reasoning
por: Kempt, Hendrik, et al.
Publicado: (2026)
por: Kempt, Hendrik, et al.
Publicado: (2026)
R^3-VQA: "Read the Room" by Video Social Reasoning
por: Niu, Lixing, et al.
Publicado: (2025)
por: Niu, Lixing, et al.
Publicado: (2025)
SafeCoT: Improving VLM Safety with Minimal Reasoning
por: Ma, Jiachen, et al.
Publicado: (2025)
por: Ma, Jiachen, et al.
Publicado: (2025)
MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents
por: Zhang, Ruoxuan, et al.
Publicado: (2025)
por: Zhang, Ruoxuan, et al.
Publicado: (2025)
SD-E$^2$: Semantic Exploration for Reasoning Under Token Budgets
por: Mishra, Kshitij, et al.
Publicado: (2026)
por: Mishra, Kshitij, et al.
Publicado: (2026)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
por: Shrestha, Robik, et al.
Publicado: (2020)
por: Shrestha, Robik, et al.
Publicado: (2020)
Ejemplares similares
-
SyntaxShap: Syntax-aware Explainability Method for Text Generation
por: Amara, Kenza, et al.
Publicado: (2024) -
Challenges and Opportunities in Text Generation Explainability
por: Amara, Kenza, et al.
Publicado: (2024) -
Concept-Level Explainability for Auditing & Steering LLM Responses
por: Amara, Kenza, et al.
Publicado: (2025) -
Deconstructing Human-AI Collaboration: Agency, Interaction, and Adaptation
por: Holter, Steffen, et al.
Publicado: (2024) -
Navigating the Maze of Explainable AI: A Systematic Approach to Evaluating Methods and Metrics
por: Klein, Lukas, et al.
Publicado: (2024)