RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Vogel, Alexander, Moured, Omar, Chen, Yufan, Zhang, Jiaming, Stiefelhagen, Rainer
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916798459281408
author Vogel, Alexander
Moured, Omar
Chen, Yufan
Zhang, Jiaming
Stiefelhagen, Rainer
author_facet Vogel, Alexander
Moured, Omar
Chen, Yufan
Zhang, Jiaming
Stiefelhagen, Rainer
contents Recently, Vision Language Models (VLMs) have increasingly emphasized document visual grounding to achieve better human-computer interaction, accessibility, and detailed understanding. However, its application to visualizations such as charts remains under-explored due to the inherent complexity of interleaved visual-numerical relationships in chart images. Existing chart understanding methods primarily focus on answering questions without explicitly identifying the visual elements that support their predictions. To bridge this gap, we introduce RefChartQA, a novel benchmark that integrates Chart Question Answering (ChartQA) with visual grounding, enabling models to refer elements at multiple granularities within chart images. Furthermore, we conduct a comprehensive evaluation by instruction-tuning 5 state-of-the-art VLMs across different categories. Our experiments demonstrate that incorporating spatial awareness via grounding improves response accuracy by over 15%, reducing hallucinations, and improving model reliability. Additionally, we identify key factors influencing text-spatial alignment, such as architectural improvements in TinyChart, which leverages a token-merging module for enhanced feature fusion. Our dataset is open-sourced for community development and further advancements. All models and code will be publicly available at https://github.com/moured/RefChartQA.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23131
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning
Vogel, Alexander
Moured, Omar
Chen, Yufan
Zhang, Jiaming
Stiefelhagen, Rainer
Computer Vision and Pattern Recognition
Recently, Vision Language Models (VLMs) have increasingly emphasized document visual grounding to achieve better human-computer interaction, accessibility, and detailed understanding. However, its application to visualizations such as charts remains under-explored due to the inherent complexity of interleaved visual-numerical relationships in chart images. Existing chart understanding methods primarily focus on answering questions without explicitly identifying the visual elements that support their predictions. To bridge this gap, we introduce RefChartQA, a novel benchmark that integrates Chart Question Answering (ChartQA) with visual grounding, enabling models to refer elements at multiple granularities within chart images. Furthermore, we conduct a comprehensive evaluation by instruction-tuning 5 state-of-the-art VLMs across different categories. Our experiments demonstrate that incorporating spatial awareness via grounding improves response accuracy by over 15%, reducing hallucinations, and improving model reliability. Additionally, we identify key factors influencing text-spatial alignment, such as architectural improvements in TinyChart, which leverages a token-merging module for enhanced feature fusion. Our dataset is open-sourced for community development and further advancements. All models and code will be publicly available at https://github.com/moured/RefChartQA.
title RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.23131