InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Iyengar, Anirudh Iyengar Kaniyar Narayana, Mukhopadhyay, Srija, Qidwai, Adnan, Singh, Shubhankar, Roth, Dan, Gupta, Vivek
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914522431750144
author Iyengar, Anirudh Iyengar Kaniyar Narayana
Mukhopadhyay, Srija
Qidwai, Adnan
Singh, Shubhankar
Roth, Dan
Gupta, Vivek
author_facet Iyengar, Anirudh Iyengar Kaniyar Narayana
Mukhopadhyay, Srija
Qidwai, Adnan
Singh, Shubhankar
Roth, Dan
Gupta, Vivek
contents We introduce InterChart, a diagnostic benchmark that evaluates how well vision-language models (VLMs) reason across multiple related charts, a task central to real-world applications such as scientific reporting, financial analysis, and public policy dashboards. Unlike prior benchmarks focusing on isolated, visually uniform charts, InterChart challenges models with diverse question types ranging from entity inference and trend correlation to numerical estimation and abstract multi-step reasoning grounded in 2-3 thematically or structurally related charts. We organize the benchmark into three tiers of increasing difficulty: (1) factual reasoning over individual charts, (2) integrative analysis across synthetically aligned chart sets, and (3) semantic inference over visually complex, real-world chart pairs. Our evaluation of state-of-the-art open- and closed-source VLMs reveals consistent and steep accuracy declines as chart complexity increases. We find that models perform better when we decompose multi-entity charts into simpler visual units, underscoring their struggles with cross-chart integration. By exposing these systematic limitations, InterChart provides a rigorous framework for advancing multimodal reasoning in complex, multi-visual environments.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07630
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
Iyengar, Anirudh Iyengar Kaniyar Narayana
Mukhopadhyay, Srija
Qidwai, Adnan
Singh, Shubhankar
Roth, Dan
Gupta, Vivek
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.7; I.2.10; I.4.10; I.7.5
We introduce InterChart, a diagnostic benchmark that evaluates how well vision-language models (VLMs) reason across multiple related charts, a task central to real-world applications such as scientific reporting, financial analysis, and public policy dashboards. Unlike prior benchmarks focusing on isolated, visually uniform charts, InterChart challenges models with diverse question types ranging from entity inference and trend correlation to numerical estimation and abstract multi-step reasoning grounded in 2-3 thematically or structurally related charts. We organize the benchmark into three tiers of increasing difficulty: (1) factual reasoning over individual charts, (2) integrative analysis across synthetically aligned chart sets, and (3) semantic inference over visually complex, real-world chart pairs. Our evaluation of state-of-the-art open- and closed-source VLMs reveals consistent and steep accuracy declines as chart complexity increases. We find that models perform better when we decompose multi-entity charts into simpler visual units, underscoring their struggles with cross-chart integration. By exposing these systematic limitations, InterChart provides a rigorous framework for advancing multimodal reasoning in complex, multi-visual environments.
title InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.7; I.2.10; I.4.10; I.7.5
url https://arxiv.org/abs/2508.07630