ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Zhihang, Bao, Xiaoyi, Li, Pandeng, Zhou, Junjie, Liao, Zhaohe, He, Yefei, Jiang, Kaixun, Xie, Chen-Wei, Zheng, Yun, Xie, Hongtao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908913645912064
author Liu, Zhihang
Bao, Xiaoyi
Li, Pandeng
Zhou, Junjie
Liao, Zhaohe
He, Yefei
Jiang, Kaixun
Xie, Chen-Wei
Zheng, Yun
Xie, Hongtao
author_facet Liu, Zhihang
Bao, Xiaoyi
Li, Pandeng
Zhou, Junjie
Liao, Zhaohe
He, Yefei
Jiang, Kaixun
Xie, Chen-Wei
Zheng, Yun
Xie, Hongtao
contents While existing generation and unified models excel at general image generation, they struggle with tasks requiring deep reasoning, planning, and precise data-to-visual mapping abilities beyond general scenarios. To push beyond the existing limitations, we introduce a new and challenging task: creative table visualization, requiring the model to generate an infographic that faithfully and aesthetically visualizes the data from a given table. To address this challenge, we propose ShowTable, a pipeline that synergizes MLLMs with diffusion models via a progressive self-correcting process. The MLLM acts as the central orchestrator for reasoning the visual plan and judging visual errors to provide refined instructions, the diffusion execute the commands from MLLM, achieving high-fidelity results. To support this task and our pipeline, we introduce three automated data construction pipelines for training different modules. Furthermore, we introduce TableVisBench, a new benchmark with 800 challenging instances across 5 evaluation dimensions, to assess performance on this task. Experiments demonstrate that our pipeline, instantiated with different models, significantly outperforms baselines, highlighting its effective multi-modal reasoning, generation, and error correction capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13303
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement
Liu, Zhihang
Bao, Xiaoyi
Li, Pandeng
Zhou, Junjie
Liao, Zhaohe
He, Yefei
Jiang, Kaixun
Xie, Chen-Wei
Zheng, Yun
Xie, Hongtao
Computer Vision and Pattern Recognition
While existing generation and unified models excel at general image generation, they struggle with tasks requiring deep reasoning, planning, and precise data-to-visual mapping abilities beyond general scenarios. To push beyond the existing limitations, we introduce a new and challenging task: creative table visualization, requiring the model to generate an infographic that faithfully and aesthetically visualizes the data from a given table. To address this challenge, we propose ShowTable, a pipeline that synergizes MLLMs with diffusion models via a progressive self-correcting process. The MLLM acts as the central orchestrator for reasoning the visual plan and judging visual errors to provide refined instructions, the diffusion execute the commands from MLLM, achieving high-fidelity results. To support this task and our pipeline, we introduce three automated data construction pipelines for training different modules. Furthermore, we introduce TableVisBench, a new benchmark with 800 challenging instances across 5 evaluation dimensions, to assess performance on this task. Experiments demonstrate that our pipeline, instantiated with different models, significantly outperforms baselines, highlighting its effective multi-modal reasoning, generation, and error correction capabilities.
title ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.13303