Mind the Uncertainty in Human Disagreement: Evaluating Discrepancies between Model Predictions and Human Responses in VQA
Fuente:
arXiv
Salvato in:
| Autori principali: | Lan, Jian, Frassinelli, Diego, Plank, Barbara |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language
di: Kim, Yoonshik, et al.
Pubblicazione: (2025)
di: Kim, Yoonshik, et al.
Pubblicazione: (2025)
Improving Automatic VQA Evaluation Using Large Language Models
di: Mañas, Oscar, et al.
Pubblicazione: (2023)
di: Mañas, Oscar, et al.
Pubblicazione: (2023)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
di: Lu, Yujie, et al.
Pubblicazione: (2024)
di: Lu, Yujie, et al.
Pubblicazione: (2024)
Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA
di: Fan, Yue, et al.
Pubblicazione: (2024)
di: Fan, Yue, et al.
Pubblicazione: (2024)
Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA
di: Karim, A H M Rezaul, et al.
Pubblicazione: (2025)
di: Karim, A H M Rezaul, et al.
Pubblicazione: (2025)
GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation
di: Zong, Yi, et al.
Pubblicazione: (2024)
di: Zong, Yi, et al.
Pubblicazione: (2024)
DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
di: Sivakumaran, Nithin, et al.
Pubblicazione: (2025)
di: Sivakumaran, Nithin, et al.
Pubblicazione: (2025)
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
di: Baek, Jeonghun, et al.
Pubblicazione: (2025)
di: Baek, Jeonghun, et al.
Pubblicazione: (2025)
Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models
di: Zhu, Wenhui, et al.
Pubblicazione: (2025)
di: Zhu, Wenhui, et al.
Pubblicazione: (2025)
Knowledge Generation for Zero-shot Knowledge-based VQA
di: Cao, Rui, et al.
Pubblicazione: (2024)
di: Cao, Rui, et al.
Pubblicazione: (2024)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
di: Shrestha, Robik, et al.
Pubblicazione: (2020)
di: Shrestha, Robik, et al.
Pubblicazione: (2020)
Towards Predicting Any Human Trajectory In Context
di: Fujii, Ryo, et al.
Pubblicazione: (2025)
di: Fujii, Ryo, et al.
Pubblicazione: (2025)
V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model Interaction
di: Zhao, Yiming, et al.
Pubblicazione: (2025)
di: Zhao, Yiming, et al.
Pubblicazione: (2025)
GDCNet: Generative Discrepancy Comparison Network for Multimodal Sarcasm Detection
di: Zhang, Shuguang, et al.
Pubblicazione: (2026)
di: Zhang, Shuguang, et al.
Pubblicazione: (2026)
Caption First, VQA Second: Knowledge Density, Not Task Format, Drives Multimodal Scaling
di: Zou, Hongjian, et al.
Pubblicazione: (2026)
di: Zou, Hongjian, et al.
Pubblicazione: (2026)
Adversarial Attacks on VQA-NLE: Exposing and Alleviating Inconsistencies in Visual Question Answering Explanations
di: Yeh, Yahsin, et al.
Pubblicazione: (2025)
di: Yeh, Yahsin, et al.
Pubblicazione: (2025)
UCSF-PDGM-VQA: Visual Question Answering dataset for brain tumor MRI interpretation
di: Ghosh, Shiv, et al.
Pubblicazione: (2026)
di: Ghosh, Shiv, et al.
Pubblicazione: (2026)
Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants
di: Qin, Lixiong, et al.
Pubblicazione: (2025)
di: Qin, Lixiong, et al.
Pubblicazione: (2025)
MindCube: Spatial Mental Modeling from Limited Views
di: Wang, Qineng, et al.
Pubblicazione: (2025)
di: Wang, Qineng, et al.
Pubblicazione: (2025)
Describing Images $\textit{Fast and Slow}$: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
di: Takmaz, Ece, et al.
Pubblicazione: (2024)
di: Takmaz, Ece, et al.
Pubblicazione: (2024)
MindBench: A Comprehensive Benchmark for Mind Map Structure Recognition and Analysis
di: Chen, Lei, et al.
Pubblicazione: (2024)
di: Chen, Lei, et al.
Pubblicazione: (2024)
HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models
di: Kang, Zhaolu, et al.
Pubblicazione: (2025)
di: Kang, Zhaolu, et al.
Pubblicazione: (2025)
RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
di: Butsanets, Léo, et al.
Pubblicazione: (2025)
di: Butsanets, Léo, et al.
Pubblicazione: (2025)
VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation
di: Park, Seongheon, et al.
Pubblicazione: (2026)
di: Park, Seongheon, et al.
Pubblicazione: (2026)
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
di: Wu, Keming, et al.
Pubblicazione: (2025)
di: Wu, Keming, et al.
Pubblicazione: (2025)
MemeMind: A Large-Scale Multimodal Dataset with Chain-of-Thought Reasoning for Harmful Meme Detection
di: Gu, Hexiang, et al.
Pubblicazione: (2025)
di: Gu, Hexiang, et al.
Pubblicazione: (2025)
Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations
di: Srivastava, Archita, et al.
Pubblicazione: (2025)
di: Srivastava, Archita, et al.
Pubblicazione: (2025)
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
di: Faure, Gueter Josmy, et al.
Pubblicazione: (2026)
di: Faure, Gueter Josmy, et al.
Pubblicazione: (2026)
Unveiling the "Fairness Seesaw": Discovering and Mitigating Gender and Race Bias in Vision-Language Models
di: Lan, Jian, et al.
Pubblicazione: (2025)
di: Lan, Jian, et al.
Pubblicazione: (2025)
Draw Your Mind: Personalized Generation via Condition-Level Modeling in Text-to-Image Diffusion Models
di: Kim, Hyungjin, et al.
Pubblicazione: (2025)
di: Kim, Hyungjin, et al.
Pubblicazione: (2025)
Pixels, Patterns, but No Poetry: To See The World like Humans
di: Gao, Hongcheng, et al.
Pubblicazione: (2025)
di: Gao, Hongcheng, et al.
Pubblicazione: (2025)
Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding
di: Kawasaki, Haruka, et al.
Pubblicazione: (2026)
di: Kawasaki, Haruka, et al.
Pubblicazione: (2026)
Seeing Eye to AI: Human Alignment via Gaze-Based Response Rewards for Large Language Models
di: Lopez-Cardona, Angela, et al.
Pubblicazione: (2024)
di: Lopez-Cardona, Angela, et al.
Pubblicazione: (2024)
Q&A Prompts: Discovering Rich Visual Clues through Mining Question-Answer Prompts for VQA requiring Diverse World Knowledge
di: Wang, Haibo, et al.
Pubblicazione: (2024)
di: Wang, Haibo, et al.
Pubblicazione: (2024)
DermaVQA-DAS: Dermatology Assessment Schema (DAS) & Datasets for Closed-Ended Question Answering & Segmentation in Patient-Generated Dermatology Images
di: Yim, Wen-wai, et al.
Pubblicazione: (2025)
di: Yim, Wen-wai, et al.
Pubblicazione: (2025)
NegVQA: Can Vision Language Models Understand Negation?
di: Zhang, Yuhui, et al.
Pubblicazione: (2025)
di: Zhang, Yuhui, et al.
Pubblicazione: (2025)
How Confident are Video Models? Empowering Video Models to Express their Uncertainty
di: Mei, Zhiting, et al.
Pubblicazione: (2025)
di: Mei, Zhiting, et al.
Pubblicazione: (2025)
T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA
di: Mo, Wentao, et al.
Pubblicazione: (2024)
di: Mo, Wentao, et al.
Pubblicazione: (2024)
Enhancing Human-Computer Interaction in Chest X-ray Analysis using Vision and Language Model with Eye Gaze Patterns
di: Kim, Yunsoo, et al.
Pubblicazione: (2024)
di: Kim, Yunsoo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language
di: Kim, Yoonshik, et al.
Pubblicazione: (2025) -
Improving Automatic VQA Evaluation Using Large Language Models
di: Mañas, Oscar, et al.
Pubblicazione: (2023) -
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
di: Lu, Yujie, et al.
Pubblicazione: (2024) -
Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA
di: Fan, Yue, et al.
Pubblicazione: (2024) -
Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA
di: Karim, A H M Rezaul, et al.
Pubblicazione: (2025)