Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy
Fuente:
arXiv
Salvato in:
| Autori principali: | Ging, Simon, Bravo, María A., Brox, Thomas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Using Knowledge Graphs to harvest datasets for efficient CLIP model training
di: Ging, Simon, et al.
Pubblicazione: (2025)
di: Ging, Simon, et al.
Pubblicazione: (2025)
Learning to Read Where to Look: Disease-Aware Vision-Language Pretraining for 3D CT
di: Ging, Simon, et al.
Pubblicazione: (2026)
di: Ging, Simon, et al.
Pubblicazione: (2026)
Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models
di: Schrodi, Simon, et al.
Pubblicazione: (2024)
di: Schrodi, Simon, et al.
Pubblicazione: (2024)
Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations
di: Srivastava, Archita, et al.
Pubblicazione: (2025)
di: Srivastava, Archita, et al.
Pubblicazione: (2025)
NegVQA: Can Vision Language Models Understand Negation?
di: Zhang, Yuhui, et al.
Pubblicazione: (2025)
di: Zhang, Yuhui, et al.
Pubblicazione: (2025)
Attribute Diversity Determines the Systematicity Gap in VQA
di: Berlot-Attwell, Ian, et al.
Pubblicazione: (2023)
di: Berlot-Attwell, Ian, et al.
Pubblicazione: (2023)
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
di: Deitke, Matt, et al.
Pubblicazione: (2024)
di: Deitke, Matt, et al.
Pubblicazione: (2024)
BloomVQA: Assessing Hierarchical Multi-modal Comprehension
di: Gong, Yunye, et al.
Pubblicazione: (2023)
di: Gong, Yunye, et al.
Pubblicazione: (2023)
Improving Automatic VQA Evaluation Using Large Language Models
di: Mañas, Oscar, et al.
Pubblicazione: (2023)
di: Mañas, Oscar, et al.
Pubblicazione: (2023)
When and How Does CLIP Enable Domain and Compositional Generalization?
di: Kempf, Elias, et al.
Pubblicazione: (2025)
di: Kempf, Elias, et al.
Pubblicazione: (2025)
Concept Bottleneck Models Without Predefined Concepts
di: Schrodi, Simon, et al.
Pubblicazione: (2024)
di: Schrodi, Simon, et al.
Pubblicazione: (2024)
O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model
di: Gupta, Rishi, et al.
Pubblicazione: (2025)
di: Gupta, Rishi, et al.
Pubblicazione: (2025)
The Neglected Tails in Vision-Language Models
di: Parashar, Shubham, et al.
Pubblicazione: (2024)
di: Parashar, Shubham, et al.
Pubblicazione: (2024)
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
di: Rädsch, Tim, et al.
Pubblicazione: (2025)
di: Rädsch, Tim, et al.
Pubblicazione: (2025)
A Vision Check-up for Language Models
di: Sharma, Pratyusha, et al.
Pubblicazione: (2024)
di: Sharma, Pratyusha, et al.
Pubblicazione: (2024)
Calibrated Self-Rewarding Vision Language Models
di: Zhou, Yiyang, et al.
Pubblicazione: (2024)
di: Zhou, Yiyang, et al.
Pubblicazione: (2024)
Overconfidence is Key: Verbalized Uncertainty Evaluation in Large Language and Vision-Language Models
di: Groot, Tobias, et al.
Pubblicazione: (2024)
di: Groot, Tobias, et al.
Pubblicazione: (2024)
Kiki or Bouba? Sound Symbolism in Vision-and-Language Models
di: Alper, Morris, et al.
Pubblicazione: (2023)
di: Alper, Morris, et al.
Pubblicazione: (2023)
Matryoshka Query Transformer for Large Vision-Language Models
di: Hu, Wenbo, et al.
Pubblicazione: (2024)
di: Hu, Wenbo, et al.
Pubblicazione: (2024)
A Survey on Hallucination in Large Vision-Language Models
di: Liu, Hanchao, et al.
Pubblicazione: (2024)
di: Liu, Hanchao, et al.
Pubblicazione: (2024)
Sherlock: Self-Correcting Reasoning in Vision-Language Models
di: Ding, Yi, et al.
Pubblicazione: (2025)
di: Ding, Yi, et al.
Pubblicazione: (2025)
IPO: Interpretable Prompt Optimization for Vision-Language Models
di: Du, Yingjun, et al.
Pubblicazione: (2024)
di: Du, Yingjun, et al.
Pubblicazione: (2024)
ICONS: Influence Consensus for Vision-Language Data Selection
di: Wu, Xindi, et al.
Pubblicazione: (2024)
di: Wu, Xindi, et al.
Pubblicazione: (2024)
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
di: Li, Siting, et al.
Pubblicazione: (2024)
di: Li, Siting, et al.
Pubblicazione: (2024)
Adding simple structure at inference improves Vision-Language Compositionality
di: Miranda, Imanol, et al.
Pubblicazione: (2025)
di: Miranda, Imanol, et al.
Pubblicazione: (2025)
Efficient Vision-Language Reasoning via Adaptive Token Pruning
di: Li, Xue, et al.
Pubblicazione: (2025)
di: Li, Xue, et al.
Pubblicazione: (2025)
SOLO: A Single Transformer for Scalable Vision-Language Modeling
di: Chen, Yangyi, et al.
Pubblicazione: (2024)
di: Chen, Yangyi, et al.
Pubblicazione: (2024)
MoPD: Mixture-of-Prompts Distillation for Vision-Language Models
di: Chen, Yang, et al.
Pubblicazione: (2024)
di: Chen, Yang, et al.
Pubblicazione: (2024)
Vision-Language Models Create Cross-Modal Task Representations
di: Luo, Grace, et al.
Pubblicazione: (2024)
di: Luo, Grace, et al.
Pubblicazione: (2024)
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
di: Li, Ang, et al.
Pubblicazione: (2025)
di: Li, Ang, et al.
Pubblicazione: (2025)
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
di: Zhou, Yiyang, et al.
Pubblicazione: (2023)
di: Zhou, Yiyang, et al.
Pubblicazione: (2023)
S-GRPO: Unified Post-Training for Large Vision-Language Models
di: Yan, Yuming, et al.
Pubblicazione: (2026)
di: Yan, Yuming, et al.
Pubblicazione: (2026)
TroL: Traversal of Layers for Large Language and Vision Models
di: Lee, Byung-Kwan, et al.
Pubblicazione: (2024)
di: Lee, Byung-Kwan, et al.
Pubblicazione: (2024)
Benchmarking Vision-Language Models for French PDF-to-Markdown Conversion
di: Rigal, Bruno, et al.
Pubblicazione: (2026)
di: Rigal, Bruno, et al.
Pubblicazione: (2026)
Towards Statistical Factuality Guarantee for Large Vision-Language Models
di: Li, Zhuohang, et al.
Pubblicazione: (2025)
di: Li, Zhuohang, et al.
Pubblicazione: (2025)
Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models
di: Chen, Yangyi, et al.
Pubblicazione: (2023)
di: Chen, Yangyi, et al.
Pubblicazione: (2023)
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts
di: Singh, Shubhankar, et al.
Pubblicazione: (2024)
di: Singh, Shubhankar, et al.
Pubblicazione: (2024)
Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
di: Ratzlaff, Neale, et al.
Pubblicazione: (2024)
di: Ratzlaff, Neale, et al.
Pubblicazione: (2024)
Learning Self-Correction in Vision-Language Models via Rollout Augmentation
di: Ding, Yi, et al.
Pubblicazione: (2026)
di: Ding, Yi, et al.
Pubblicazione: (2026)
When are Lemons Purple? The Concept Association Bias of Vision-Language Models
di: Yamada, Yutaro, et al.
Pubblicazione: (2022)
di: Yamada, Yutaro, et al.
Pubblicazione: (2022)
Documenti analoghi
-
Using Knowledge Graphs to harvest datasets for efficient CLIP model training
di: Ging, Simon, et al.
Pubblicazione: (2025) -
Learning to Read Where to Look: Disease-Aware Vision-Language Pretraining for 3D CT
di: Ging, Simon, et al.
Pubblicazione: (2026) -
Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models
di: Schrodi, Simon, et al.
Pubblicazione: (2024) -
Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations
di: Srivastava, Archita, et al.
Pubblicazione: (2025) -
NegVQA: Can Vision Language Models Understand Negation?
di: Zhang, Yuhui, et al.
Pubblicazione: (2025)