ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Zhengzhuo, Du, SiNan, Qi, Yiyan, SiwenLu, Xu, Chengjin, Yuan, Chun, Guo, Jian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915644719497216
author Xu, Zhengzhuo
Du, SiNan
Qi, Yiyan
SiwenLu
Xu, Chengjin
Yuan, Chun
Guo, Jian
author_facet Xu, Zhengzhuo
Du, SiNan
Qi, Yiyan
SiwenLu
Xu, Chengjin
Yuan, Chun
Guo, Jian
contents Multimodal Large Language Models (MLLMs) have emerged as powerful tools for chart comprehension. However, they heavily rely on extracted content via OCR, which leads to numerical hallucinations when chart textual annotations are sparse. While existing methods focus on scaling instructions, they fail to address the fundamental challenge, i.e., reasoning with visual perception. In this paper, we identify a critical observation: MLLMs exhibit weak grounding in chart elements and proportional relationships, as evidenced by their inability to localize key positions to match their reasoning. To bridge this gap, we propose PointCoT, which integrates reflective interaction into chain-of-thought reasoning in charts. By prompting MLLMs to generate bounding boxes and re-render charts based on location annotations, we establish connections between textual reasoning steps and visual grounding regions. We further introduce an automated pipeline to construct ChartPoint-SFT-62k, a dataset featuring 19.2K high-quality chart samples with step-by-step CoT, bounding box, and re-rendered visualizations. Leveraging this data, we develop two instruction-tuned models, ChartPointQ2 and ChartPointQ2.5, which outperform state-of-the-art across several chart benchmarks, e.g., +5.04\% on ChartBench.
format Preprint
id arxiv_https___arxiv_org_abs_2512_00305
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
Xu, Zhengzhuo
Du, SiNan
Qi, Yiyan
SiwenLu
Xu, Chengjin
Yuan, Chun
Guo, Jian
Artificial Intelligence
Multimodal Large Language Models (MLLMs) have emerged as powerful tools for chart comprehension. However, they heavily rely on extracted content via OCR, which leads to numerical hallucinations when chart textual annotations are sparse. While existing methods focus on scaling instructions, they fail to address the fundamental challenge, i.e., reasoning with visual perception. In this paper, we identify a critical observation: MLLMs exhibit weak grounding in chart elements and proportional relationships, as evidenced by their inability to localize key positions to match their reasoning. To bridge this gap, we propose PointCoT, which integrates reflective interaction into chain-of-thought reasoning in charts. By prompting MLLMs to generate bounding boxes and re-render charts based on location annotations, we establish connections between textual reasoning steps and visual grounding regions. We further introduce an automated pipeline to construct ChartPoint-SFT-62k, a dataset featuring 19.2K high-quality chart samples with step-by-step CoT, bounding box, and re-rendered visualizations. Leveraging this data, we develop two instruction-tuned models, ChartPointQ2 and ChartPointQ2.5, which outperform state-of-the-art across several chart benchmarks, e.g., +5.04\% on ChartBench.
title ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
topic Artificial Intelligence
url https://arxiv.org/abs/2512.00305