SimVecVis: A Dataset for Enhancing MLLMs in Visualization Understanding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Can, Da, Chunlin, Long, Xiaoxiao, Yang, Yuxiao, Zhang, Yu, Wang, Yong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916820726841344
author Liu, Can
Da, Chunlin
Long, Xiaoxiao
Yang, Yuxiao
Zhang, Yu
Wang, Yong
author_facet Liu, Can
Da, Chunlin
Long, Xiaoxiao
Yang, Yuxiao
Zhang, Yu
Wang, Yong
contents Current multimodal large language models (MLLMs), while effective in natural image understanding, struggle with visualization understanding due to their inability to decode the data-to-visual mapping and extract structured information. To address these challenges, we propose SimVec, a novel simplified vector format that encodes chart elements such as mark type, position, and size. The effectiveness of SimVec is demonstrated by using MLLMs to reconstruct chart information from SimVec formats. Then, we build a new visualization dataset, SimVecVis, to enhance the performance of MLLMs in visualization understanding, which consists of three key dimensions: bitmap images of charts, their SimVec representations, and corresponding data-centric question-answering (QA) pairs with explanatory chain-of-thought (CoT) descriptions. We finetune state-of-the-art MLLMs (e.g., MiniCPM and Qwen-VL), using SimVecVis with different dataset dimensions. The experimental results show that it leads to substantial performance improvements of MLLMs with good spatial perception capabilities (e.g., MiniCPM) in data-centric QA tasks. Our dataset and source code are available at: https://github.com/VIDA-Lab/SimVecVis.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21319
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SimVecVis: A Dataset for Enhancing MLLMs in Visualization Understanding
Liu, Can
Da, Chunlin
Long, Xiaoxiao
Yang, Yuxiao
Zhang, Yu
Wang, Yong
Human-Computer Interaction
Computer Vision and Pattern Recognition
Current multimodal large language models (MLLMs), while effective in natural image understanding, struggle with visualization understanding due to their inability to decode the data-to-visual mapping and extract structured information. To address these challenges, we propose SimVec, a novel simplified vector format that encodes chart elements such as mark type, position, and size. The effectiveness of SimVec is demonstrated by using MLLMs to reconstruct chart information from SimVec formats. Then, we build a new visualization dataset, SimVecVis, to enhance the performance of MLLMs in visualization understanding, which consists of three key dimensions: bitmap images of charts, their SimVec representations, and corresponding data-centric question-answering (QA) pairs with explanatory chain-of-thought (CoT) descriptions. We finetune state-of-the-art MLLMs (e.g., MiniCPM and Qwen-VL), using SimVecVis with different dataset dimensions. The experimental results show that it leads to substantial performance improvements of MLLMs with good spatial perception capabilities (e.g., MiniCPM) in data-centric QA tasks. Our dataset and source code are available at: https://github.com/VIDA-Lab/SimVecVis.
title SimVecVis: A Dataset for Enhancing MLLMs in Visualization Understanding
topic Human-Computer Interaction
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.21319