WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Foroutan, Negar, Romanou, Angelika, Ansaripour, Matin, Eisenschlos, Julian Martin, Aberer, Karl, Lebret, Rémi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911012149526528
author Foroutan, Negar
Romanou, Angelika
Ansaripour, Matin
Eisenschlos, Julian Martin
Aberer, Karl
Lebret, Rémi
author_facet Foroutan, Negar
Romanou, Angelika
Ansaripour, Matin
Eisenschlos, Julian Martin
Aberer, Karl
Lebret, Rémi
contents Documents are fundamental to preserving and disseminating information, often incorporating complex layouts, tables, and charts that pose significant challenges for automatic document understanding (DU). While vision-language large models (VLLMs) have demonstrated improvements across various tasks, their effectiveness in processing long-context vision inputs remains unclear. This paper introduces WikiMixQA, a benchmark comprising 1,000 multiple-choice questions (MCQs) designed to evaluate cross-modal reasoning over tables and charts extracted from 4,000 Wikipedia pages spanning seven distinct topics. Unlike existing benchmarks, WikiMixQA emphasizes complex reasoning by requiring models to synthesize information from multiple modalities. We evaluate 12 state-of-the-art vision-language models, revealing that while proprietary models achieve ~70% accuracy when provided with direct context, their performance deteriorates significantly when retrieval from long documents is required. Among these, GPT-4-o is the only model exceeding 50% accuracy in this setting, whereas open-source models perform considerably worse, with a maximum accuracy of 27%. These findings underscore the challenges of long-context, multi-modal reasoning and establish WikiMixQA as a crucial benchmark for advancing document understanding research.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15594
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
Foroutan, Negar
Romanou, Angelika
Ansaripour, Matin
Eisenschlos, Julian Martin
Aberer, Karl
Lebret, Rémi
Computation and Language
Artificial Intelligence
Machine Learning
Documents are fundamental to preserving and disseminating information, often incorporating complex layouts, tables, and charts that pose significant challenges for automatic document understanding (DU). While vision-language large models (VLLMs) have demonstrated improvements across various tasks, their effectiveness in processing long-context vision inputs remains unclear. This paper introduces WikiMixQA, a benchmark comprising 1,000 multiple-choice questions (MCQs) designed to evaluate cross-modal reasoning over tables and charts extracted from 4,000 Wikipedia pages spanning seven distinct topics. Unlike existing benchmarks, WikiMixQA emphasizes complex reasoning by requiring models to synthesize information from multiple modalities. We evaluate 12 state-of-the-art vision-language models, revealing that while proprietary models achieve ~70% accuracy when provided with direct context, their performance deteriorates significantly when retrieval from long documents is required. Among these, GPT-4-o is the only model exceeding 50% accuracy in this setting, whereas open-source models perform considerably worse, with a maximum accuracy of 27%. These findings underscore the challenges of long-context, multi-modal reasoning and establish WikiMixQA as a crucial benchmark for advancing document understanding research.
title WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.15594