ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Xuanle, Luo, Xianzhen, Shi, Qi, Chen, Chi, Wang, Shuo, Liu, Zhiyuan, Sun, Maosong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913921701511168
author Zhao, Xuanle
Luo, Xianzhen
Shi, Qi
Chen, Chi
Wang, Shuo
Liu, Zhiyuan
Sun, Maosong
author_facet Zhao, Xuanle
Luo, Xianzhen
Shi, Qi
Chen, Chi
Wang, Shuo
Liu, Zhiyuan
Sun, Maosong
contents Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in chart understanding tasks. However, interpreting charts with textual descriptions often leads to information loss, as it fails to fully capture the dense information embedded in charts. In contrast, parsing charts into code provides lossless representations that can effectively contain all critical details. Although existing open-source MLLMs have achieved success in chart understanding tasks, they still face two major challenges when applied to chart-to-code tasks: (1) Low executability and poor restoration of chart details in the generated code and (2) Lack of large-scale and diverse training data. To address these challenges, we propose \textbf{ChartCoder}, the first dedicated chart-to-code MLLM, which leverages Code LLMs as the language backbone to enhance the executability of the generated code. Furthermore, we introduce \textbf{Chart2Code-160k}, the first large-scale and diverse dataset for chart-to-code generation, and propose the \textbf{Snippet-of-Thought (SoT)} method, which transforms direct chart-to-code generation data into step-by-step generation. Experiments demonstrate that ChartCoder, with only 7B parameters, surpasses existing open-source MLLMs on chart-to-code benchmarks, achieving superior chart restoration and code excitability. Our code is available at https://github.com/thunlp/ChartCoder.
format Preprint
id arxiv_https___arxiv_org_abs_2501_06598
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation
Zhao, Xuanle
Luo, Xianzhen
Shi, Qi
Chen, Chi
Wang, Shuo
Liu, Zhiyuan
Sun, Maosong
Artificial Intelligence
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in chart understanding tasks. However, interpreting charts with textual descriptions often leads to information loss, as it fails to fully capture the dense information embedded in charts. In contrast, parsing charts into code provides lossless representations that can effectively contain all critical details. Although existing open-source MLLMs have achieved success in chart understanding tasks, they still face two major challenges when applied to chart-to-code tasks: (1) Low executability and poor restoration of chart details in the generated code and (2) Lack of large-scale and diverse training data. To address these challenges, we propose \textbf{ChartCoder}, the first dedicated chart-to-code MLLM, which leverages Code LLMs as the language backbone to enhance the executability of the generated code. Furthermore, we introduce \textbf{Chart2Code-160k}, the first large-scale and diverse dataset for chart-to-code generation, and propose the \textbf{Snippet-of-Thought (SoT)} method, which transforms direct chart-to-code generation data into step-by-step generation. Experiments demonstrate that ChartCoder, with only 7B parameters, surpasses existing open-source MLLMs on chart-to-code benchmarks, achieving superior chart restoration and code excitability. Our code is available at https://github.com/thunlp/ChartCoder.
title ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation
topic Artificial Intelligence
url https://arxiv.org/abs/2501.06598