Questioning the Stability of Visual Question Answering
Fuente:
arXiv
Salvato in:
| Autori principali: | Rosenfeld, Amir, Glazer, Neta, Fetaya, Ethan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Can Modern Vision Models Understand the Difference Between an Object and a Look-alike?
di: Cohen, Itay, et al.
Pubblicazione: (2025)
di: Cohen, Itay, et al.
Pubblicazione: (2025)
BERT-VQA: Visual Question Answering on Plots
di: Vu, Tai, et al.
Pubblicazione: (2025)
di: Vu, Tai, et al.
Pubblicazione: (2025)
Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation Models
di: Mamaghan, Amir Mohammad Karimi, et al.
Pubblicazione: (2024)
di: Mamaghan, Amir Mohammad Karimi, et al.
Pubblicazione: (2024)
Inverse Problem Sampling in Latent Space Using Sequential Monte Carlo
di: Achituve, Idan, et al.
Pubblicazione: (2025)
di: Achituve, Idan, et al.
Pubblicazione: (2025)
Privacy-Aware Document Visual Question Answering
di: Tito, Rubèn, et al.
Pubblicazione: (2023)
di: Tito, Rubèn, et al.
Pubblicazione: (2023)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
di: Souibgui, Mohamed Ali, et al.
Pubblicazione: (2025)
di: Souibgui, Mohamed Ali, et al.
Pubblicazione: (2025)
Describe Anything Model for Visual Question Answering on Text-rich Images
di: Vu, Yen-Linh, et al.
Pubblicazione: (2025)
di: Vu, Yen-Linh, et al.
Pubblicazione: (2025)
RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering
di: Wang, Yuduo, et al.
Pubblicazione: (2023)
di: Wang, Yuduo, et al.
Pubblicazione: (2023)
Learning to Compress Contexts for Efficient Knowledge-based Visual Question Answering
di: Weng, Weixi, et al.
Pubblicazione: (2024)
di: Weng, Weixi, et al.
Pubblicazione: (2024)
TPCL: Task Progressive Curriculum Learning for Robust Visual Question Answering
di: Akl, Ahmed, et al.
Pubblicazione: (2024)
di: Akl, Ahmed, et al.
Pubblicazione: (2024)
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
di: Chintapatla, Ishant, et al.
Pubblicazione: (2025)
di: Chintapatla, Ishant, et al.
Pubblicazione: (2025)
Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions
di: Indrehus, Kjetil, et al.
Pubblicazione: (2026)
di: Indrehus, Kjetil, et al.
Pubblicazione: (2026)
Exploring Diverse Methods in Visual Question Answering
di: Li, Panfeng, et al.
Pubblicazione: (2024)
di: Li, Panfeng, et al.
Pubblicazione: (2024)
Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
di: Hartsock, Iryna, et al.
Pubblicazione: (2024)
di: Hartsock, Iryna, et al.
Pubblicazione: (2024)
From Segments to Concepts: Interpretable Image Classification via Concept-Guided Segmentation
di: Eisenberg, Ran, et al.
Pubblicazione: (2025)
di: Eisenberg, Ran, et al.
Pubblicazione: (2025)
RECODE: Reasoning Through Code Generation for Visual Question Answering
di: Shen, Junhong, et al.
Pubblicazione: (2025)
di: Shen, Junhong, et al.
Pubblicazione: (2025)
SViQA: A Unified Speech-Vision Multimodal Model for Textless Visual Question Answering
di: Li, Bingxin
Pubblicazione: (2025)
di: Li, Bingxin
Pubblicazione: (2025)
ViInfographicVQA: A Benchmark for Single and Multi-image Visual Question Answering on Vietnamese Infographics
di: Van-Dinh, Tue-Thu, et al.
Pubblicazione: (2025)
di: Van-Dinh, Tue-Thu, et al.
Pubblicazione: (2025)
Evaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering Benchmarks
di: Singh, Simranjit, et al.
Pubblicazione: (2024)
di: Singh, Simranjit, et al.
Pubblicazione: (2024)
Enhancing Visual Question Answering through Ranking-Based Hybrid Training and Multimodal Fusion
di: Chen, Peiyuan, et al.
Pubblicazione: (2024)
di: Chen, Peiyuan, et al.
Pubblicazione: (2024)
CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
di: Romero, David, et al.
Pubblicazione: (2024)
di: Romero, David, et al.
Pubblicazione: (2024)
Find The Gap: Knowledge Base Reasoning For Visual Question Answering
di: Barezi, Elham J., et al.
Pubblicazione: (2024)
di: Barezi, Elham J., et al.
Pubblicazione: (2024)
Federated Document Visual Question Answering: A Pilot Study
di: Nguyen, Khanh, et al.
Pubblicazione: (2024)
di: Nguyen, Khanh, et al.
Pubblicazione: (2024)
Expert Knowledge-Aware Image Difference Graph Representation Learning for Difference-Aware Medical Visual Question Answering
di: Hu, Xinyue, et al.
Pubblicazione: (2023)
di: Hu, Xinyue, et al.
Pubblicazione: (2023)
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
di: Li, Xu, et al.
Pubblicazione: (2025)
di: Li, Xu, et al.
Pubblicazione: (2025)
Enhancing Multi-Image Question Answering via Submodular Subset Selection
di: Sharma, Aaryan, et al.
Pubblicazione: (2025)
di: Sharma, Aaryan, et al.
Pubblicazione: (2025)
Improving Video Question Answering through query-based frame selection
di: Patil, Himanshu, et al.
Pubblicazione: (2026)
di: Patil, Himanshu, et al.
Pubblicazione: (2026)
Answering Questions in Stages: Prompt Chaining for Contract QA
di: Roegiest, Adam, et al.
Pubblicazione: (2024)
di: Roegiest, Adam, et al.
Pubblicazione: (2024)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
di: Gupta, Akash, et al.
Pubblicazione: (2025)
di: Gupta, Akash, et al.
Pubblicazione: (2025)
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts
di: Singh, Shubhankar, et al.
Pubblicazione: (2024)
di: Singh, Shubhankar, et al.
Pubblicazione: (2024)
Towards Fine-Grained Video Question Answering
di: Dai, Wei, et al.
Pubblicazione: (2025)
di: Dai, Wei, et al.
Pubblicazione: (2025)
ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering
di: Diao, Xingjian, et al.
Pubblicazione: (2025)
di: Diao, Xingjian, et al.
Pubblicazione: (2025)
Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question Answering
di: Yu, Zhou, et al.
Pubblicazione: (2023)
di: Yu, Zhou, et al.
Pubblicazione: (2023)
ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response
di: Xie, Stephan, et al.
Pubblicazione: (2026)
di: Xie, Stephan, et al.
Pubblicazione: (2026)
CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering
di: Huai, Tianyu, et al.
Pubblicazione: (2025)
di: Huai, Tianyu, et al.
Pubblicazione: (2025)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
di: Lagos, Maximiliano Hormazábal, et al.
Pubblicazione: (2025)
di: Lagos, Maximiliano Hormazábal, et al.
Pubblicazione: (2025)
ReasonDrive: Efficient Visual Question Answering for Autonomous Vehicles with Reasoning-Enhanced Small Vision-Language Models
di: Chahe, Amirhosein, et al.
Pubblicazione: (2025)
di: Chahe, Amirhosein, et al.
Pubblicazione: (2025)
ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?
di: Meshram, Pragati Shuddhodhan, et al.
Pubblicazione: (2024)
di: Meshram, Pragati Shuddhodhan, et al.
Pubblicazione: (2024)
CinePile: A Long Video Question Answering Dataset and Benchmark
di: Rawal, Ruchit, et al.
Pubblicazione: (2024)
di: Rawal, Ruchit, et al.
Pubblicazione: (2024)
Exploring Advanced Techniques for Visual Question Answering: A Comprehensive Comparison
di: Baby, Aiswarya, et al.
Pubblicazione: (2025)
di: Baby, Aiswarya, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Can Modern Vision Models Understand the Difference Between an Object and a Look-alike?
di: Cohen, Itay, et al.
Pubblicazione: (2025) -
BERT-VQA: Visual Question Answering on Plots
di: Vu, Tai, et al.
Pubblicazione: (2025) -
Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation Models
di: Mamaghan, Amir Mohammad Karimi, et al.
Pubblicazione: (2024) -
Inverse Problem Sampling in Latent Space Using Sequential Monte Carlo
di: Achituve, Idan, et al.
Pubblicazione: (2025) -
Privacy-Aware Document Visual Question Answering
di: Tito, Rubèn, et al.
Pubblicazione: (2023)