Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cui, Zhiqing, Yuan, Jiahao, Wang, Hanqing, Li, Yanshu, Du, Chenxu, Ding, Zhenglong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914088824602624
author Cui, Zhiqing
Yuan, Jiahao
Wang, Hanqing
Li, Yanshu
Du, Chenxu
Ding, Zhenglong
author_facet Cui, Zhiqing
Yuan, Jiahao
Wang, Hanqing
Li, Yanshu
Du, Chenxu
Ding, Zhenglong
contents Scientific diagrams are vital tools for communicating structured knowledge across disciplines. However, they are often published as static raster images, losing symbolic semantics and limiting reuse. While Multimodal Large Language Models (MLLMs) offer a pathway to bridging vision and structure, existing methods lack semantic control and structural interpretability, especially on complex diagrams. We propose Draw with Thought (DwT), a training-free framework that guides MLLMs to reconstruct diagrams into editable mxGraph XML code through cognitively-grounded Chain-of-Thought reasoning. DwT enables interpretable and controllable outputs without model fine-tuning by dividing the task into two stages: Coarse-to-Fine Planning, which handles perceptual structuring and semantic specification, and Structure-Aware Code Generation, enhanced by format-guided refinement. To support evaluation, we release Plot2XML, a benchmark of 247 real-world scientific diagrams with gold-standard XML annotations. Extensive experiments across eight MLLMs show that our approach yields high-fidelity, semantically aligned, and structurally valid reconstructions, with human evaluations confirming strong alignment in both accuracy and visual aesthetics, offering a scalable solution for converting static visuals into executable representations and advancing machine understanding of scientific graphics.
format Preprint
id arxiv_https___arxiv_org_abs_2504_09479
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation
Cui, Zhiqing
Yuan, Jiahao
Wang, Hanqing
Li, Yanshu
Du, Chenxu
Ding, Zhenglong
Artificial Intelligence
Computation and Language
Scientific diagrams are vital tools for communicating structured knowledge across disciplines. However, they are often published as static raster images, losing symbolic semantics and limiting reuse. While Multimodal Large Language Models (MLLMs) offer a pathway to bridging vision and structure, existing methods lack semantic control and structural interpretability, especially on complex diagrams. We propose Draw with Thought (DwT), a training-free framework that guides MLLMs to reconstruct diagrams into editable mxGraph XML code through cognitively-grounded Chain-of-Thought reasoning. DwT enables interpretable and controllable outputs without model fine-tuning by dividing the task into two stages: Coarse-to-Fine Planning, which handles perceptual structuring and semantic specification, and Structure-Aware Code Generation, enhanced by format-guided refinement. To support evaluation, we release Plot2XML, a benchmark of 247 real-world scientific diagrams with gold-standard XML annotations. Extensive experiments across eight MLLMs show that our approach yields high-fidelity, semantically aligned, and structurally valid reconstructions, with human evaluations confirming strong alignment in both accuracy and visual aesthetics, offering a scalable solution for converting static visuals into executable representations and advancing machine understanding of scientific graphics.
title Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2504.09479